South Minneapolis News

collapse
Home / Daily News Analysis / I’ve been testing AI content detectors for years – these are your best options in 2025

I’ve been testing AI content detectors for years – these are your best options in 2025

Sep 07, 2026  Twila Rosenbaum  6 views
I’ve been testing AI content detectors for years – these are your best options in 2025

How hard is it in 2025, three years after generative AI tools entered everyday work, to distinguish human writing from machine output? Another round of tests shows that some dedicated AI detectors still perform well, but the category is unstable. A tool that scored 100 percent in one cycle can fail badly a few months later.

Key facts from the latest test

  • Eleven dedicated AI content detectors were tested using five writing samples, two written by humans and three written by ChatGPT.
  • Only three dedicated detectors earned perfect scores: Pangram, QuillBot, and ZeroGPT.
  • Several major detectors fell sharply in accuracy, including Undetectable.ai, which dropped from a perfect score to 20 percent since the previous test cycle.
  • Three chatbots, ChatGPT Plus, Copilot, and Gemini, also returned perfect scores, showing that a general chatbot can compete with specialized tools.
  • The tests show that AI content detectors are not reliable enough to be the sole judge of academic or editorial integrity.

Why this test matters

Generative AI is now a normal part of writing for many students, marketers, and business teams. That has made plagiarism harder to define. If a person uses an AI writing tool and passes the output off as their own work without credit, most definitions of plagiarism apply. The words may not be stolen from another author, but claiming an AI produced them under a human byline is still a form of academic and professional misconduct.

Detecting that misconduct is not simple. AI language models are trained to imitate human patterns, and those patterns change with every new model. A detector that catches older AI text may miss newer output. The same detector may also flag a clearly written human article as machine made, especially if the writer is not a native English speaker or if the text is especially clean and structured.

How the testing was done

The test used five blocks of text. Two were written by a human editor; three were generated by ChatGPT. Each block was submitted separately to each AI detector, and the result was marked either correct or incorrect. When a detector gave a percentage, anything above 70 percent confidence was treated as its answer. If a detector did not reach that confidence, the result was treated as uncertain.

This same framework has been used for several test cycles since early 2023. In the first round, the best detector scored only 66 percent. In early 2025, three of ten detectors earned perfect scores in one round and five in another. In the latest round, only three of eleven detectors were perfect, despite the number of tools growing. Two detectors that had previously been perfect dropped in quality, in some cases at the same time they added stricter limits on free use.

Content detector performance

The latest run included BrandWell, Copyleaks, GPT-2 Output Detector, GPTZero, Grammarly, Originality.ai, Pangram, QuillBot, Undetectable.ai, Writer.com, and ZeroGPT. One previously included detector, Monica, was dropped because it limited tests to 250 words and then moved its full testing tools behind a premium subscription. Pangram was added as a new entrant and immediately joined the top group.

  • BrandWell AI Content Detection: 40 percent accuracy
  • Copyleaks: 80 percent accuracy
  • GPT-2 Output Detector: 60 percent accuracy
  • GPTZero: 80 percent accuracy
  • Grammarly: 40 percent accuracy
  • Pangram: 100 percent accuracy
  • Originality.ai: 80 percent accuracy
  • QuillBot: 100 percent accuracy
  • Undetectable.ai: 20 percent accuracy
  • Writer.com AI Content Detector: 40 percent accuracy
  • ZeroGPT: 100 percent accuracy

The perfect scorers in the detector group

Pangram was the strongest new entry. It is a relative newcomer with a team that includes former engineers from Google and Tesla, and it focuses specifically on detecting AI content rather than on rewriting or humanizing text. Pangram allows a limited number of free scans per day, which was enough for this test. Processing was slower than some competing tools, but the results were accurate across all five samples.

QuillBot has had a rocky history in this test. Early rounds produced wildly inconsistent scores, with repeated scans of the same text sometimes giving conflicting answers. That changed in the most recent cycles. QuillBot returned perfect scores in consecutive tests and correctly identified every human-written and AI-written block.

ZeroGPT also continued its strong performance. Earlier versions of the site looked unfinished and were filled with advertising, but the service has matured into a more professional product with clear pricing, company information, and customer support. Its accuracy improved to 100 percent in the summer and held steady in this latest run.

Mixed performers and notable failures

Copyleaks, which has publicly claimed to be the most accurate AI detector, scored 80 percent. It correctly identified most of the AI blocks but flagged a human-written block as 100 percent AI. That kind of false positive is dangerous in academic or editorial settings because it can lead to punishment for an innocent writer.

GPTZero also scored 80 percent. It has evolved from what once looked like an experimental project into a professional service with a clear mission around protecting human writing. The tool has been updated over time, but the changes did not consistently improve accuracy. In one recent round it incorrectly identified human text and correctly caught an AI block. In this round it caught the human text and missed an AI block. The shifting pattern makes it less reliable than its overall score suggests.

Originality.ai markets itself as a highly accurate commercial detector and sells credits for scans. It scored 80 percent in this test, but it also made the significant mistake of declaring a human-written block to be 100 percent AI. That was a change from earlier rounds, when it correctly identified the same style of human writing.

Undetectable.ai took the biggest drop. It had earned a perfect score in an earlier round, but in this round it rated human writing as only 60 percent likely to be AI, and it rated three AI-written samples as likely human. That gives the tool a 20 percent accuracy score, which is effectively no better than guessing.

BrandWell, which evolved from an AI content generation service, stayed at 40 percent. It was confused by at least one ChatGPT-written block and called two other AI blocks human. Grammarly, despite its strong reputation for analyzing text, also remained at 40 percent with no meaningful improvement. Writer.com wrongly identified every block in this test as human-written, even though three of the five came from ChatGPT.

The GPT-2 Output Detector also remained stuck at 60 percent. It was built with an older language model and has apparently not been updated to handle newer AI systems. It may still be useful as a classroom curiosity, but it is not a serious option for verifying modern AI-generated text.

Can chatbots replace content detectors?

Because dedicated content detectors produced inconsistent results, this test also tried a different approach: asking ordinary AI chatbots to evaluate the same five text blocks. Each chatbot was given a simple prompt asking whether the text was written by a human or by an AI. No additional training or special configuration was used.

ChatGPT Plus returned a perfect score. It correctly identified every human-written block and every AI-written block. Copilot and Gemini also returned perfect scores. That puts the top chatbots on the same level as the best dedicated detectors, and in some cases better.

ChatGPT’s free tier also performed well, missing only one human-written block. In one of its correct answers, it went further and identified the writer of a human sample based on publicly available information. That was an unexpected result, but it illustrates how much context a general chatbot can bring to an analysis.

Grok did not perform as well. Despite strong showings in general AI assistant tests, it failed three of the five text samples and classified the majority of blocks, including AI-generated ones, as human. This suggests that chatbot-based detection is possible only with models that are especially strong at analyzing tone, context, and style.

Why chatbot results are important

The main advantage of chatbot-based detection is cost and convenience. Many people already subscribe to a chatbot service or use a free tier. They do not need to buy another monthly service just to check a suspect paragraph. The results also show that a well-designed general model can often do what specialized detectors claim to do.

That does not mean chatbot judgment is perfect. ChatGPT’s free tier made at least one error in this round. A chatbot can also be too confident when explaining its reasoning, and users may not know how much training data shaped its opinion. Still, for many basic checks, a strong chatbot is a credible alternative.

Practical caution for readers

No AI content detector should be used alone to punish a writer. Human-written text is often flagged incorrectly, and AI-generated text can be rewritten just enough to avoid the most common detection patterns. Non-native English writers have consistently been at higher risk of being falsely accused, because their phrasing can seem unusually regular to a detection algorithm.

The strongest approach is to combine tools. Run a suspicious text through one of the three perfect-scoring detectors, then ask a capable chatbot for a second opinion. Editors and teachers should still read the text carefully and give the writer a chance to explain their process. AI detection is best treated as a warning signal, not as a final verdict.

Have you tried AI content detectors this year? Have any of them wrongly flagged your writing as AI, or missed AI content that seemed obvious to you? Your own results may be the best guide to which tool you can trust.


Source: ZDNET News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy