A new study from research organization Epoch AI reveals that commercial AI text detectors, while nearly flawless at catching text generated from simple prompts, fail to identify a significant portion of AI-generated content when language models are instructed to mimic a specific author's writing style. The problem is most acute in scientific writing, the very domain where detection tools are most likely to be deployed.
The study, published by THE DECODER on July 19, 2026, tested three widely used detectors: Pangram (version 3.3.2), GPTZero (model 2026-05-11-base), and Originality.ai (Turbo 3.0.2). Researchers built a corpus of 495 human passages from 99 authors, evenly split across blogging, fiction, and scientific writing. All human texts were written before ChatGPT's release in November 2022.
Near-Perfect Results on Plain AI Text
When detectors faced plain AI-generated text—produced by standard prompts without stylistic constraints—their performance was stellar. The false-negative rate, the proportion of AI texts that slipped through as human, topped out at just 0.7% across all three tools.
Pangram and GPTZero produced zero false alarms on the human texts. Originality.ai flagged 19 out of 495 human passages as AI-generated, a false-positive rate of 3.8%.
An earlier, unspecified test by the Authors Guild had found that Pangram and Originality.ai reliably classified human texts as human. The Epoch AI study confirms that low false-alarm rate, but it shows that this says little about how many AI texts actually slip through.
Style Imitation Drives Detection Failures
The study's core finding emerged when researchers tested style imitation. Three frontier language models—Claude Opus 4.8, GPT-5.5, and Gemini 3.1 Pro—each received five real text passages from an author and wrote new text in the same style. This generated 297 style-imitated passages.
On average, 38 out of 297 style-imitated passages went undetected, yielding a false-negative rate of roughly 13%. The results varied by detector. Pangram missed 10% of style-imitated texts. GPTZero missed 11%. Originality.ai missed 18%.
Despite different underlying methods, all three detectors showed the same pattern: catch simple prompts almost every time, but miss imitations far more often. Pangram uses a neural network trained on human and machine-generated text; its founder called the system a black box since its verdicts can't be traced. GPTZero measures how predictable word choices are and the variation within text. Originality.ai searches for statistical patterns from training on human and AI-generated text.
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
Scientific Writing Is the Weakest Spot
The failures were not evenly distributed across genres. For fiction, the false-negative rate across all detectors was between 1% and 5%. Scientific writing was where detectors failed the hardest.
In that domain, Pangram missed 25% of style-imitated academic AI texts. GPTZero missed 24%. Originality.ai missed 29%. The worst individual result was Pangram missing 48% of Gemini-generated academic passages. Originality.ai missed 39% of academic GPT-5.5 texts.
This is especially troubling because scientific writing is where AI detection probably sees the most real-world use, in settings such as peer review, grant evaluation, and academic integrity checks. The study's corpus of 495 human passages, evenly split across blogging, fiction, and scientific writing, was designed to reflect that real-world diversity.
What the Findings Mean
The Epoch AI study underscores a fundamental limitation of current detection technology. A tool that rarely flags human writing as AI-generated can still let a substantial fraction of AI-generated text pass through, especially when that text is crafted to resemble a known author's voice.
The three detectors tested use different approaches, yet they converge on the same vulnerability. The Pangram founder acknowledged the opacity of his system, calling it a black box. GPTZero's method relies on statistical measures of word predictability and variation. Originality.ai trains on large corpora of human and AI-generated text to find statistical patterns. None of these approaches proved robust against style imitation in scientific writing.
The study's results suggest that users of AI text detectors—educators, publishers, and researchers—should not assume that a low false-alarm rate on human writing guarantees strong detection of AI-generated content. The gap between the 0.7% false-negative rate on plain AI text and the roughly 13% average on style-imitated text is a warning that the arms race between generation and detection is far from over.

