Research

AI Writing on arXiv Measured: One Third of New Papers Flagged

A new study scored the full text of 12,750 arXiv papers and found that roughly one third of recent submissions read as machine-written. The detector was calibrated to a 0.4% false-positive rate on pre-ChatGPT papers. Computer science leads at 65%, while mathematics is lowest at 0.7%, though the authors note limitations in detecting AI writing in notation-heavy fields.

Neura News

Neura News

Neura Market Editorial

July 20, 20265 min read

Originally reported by unslop.run

AI Writing on arXiv Measured: One Third of New Papers Flagged

Researchers have analyzed the full text of 12,750 arXiv papers and concluded that approximately one third of new submissions appear to be machine-written. The study, which covers papers from January 2023 to July 2026, provides a detailed look at how AI writing has spread across academic fields since the release of ChatGPT.

Methodology and False-Positive Floor

The researchers built their study around a common criticism of AI detectors: that they often flag genuine human writing as machine-generated. To address this, they calibrated their detector specifically for academic writing. At a 0.4% false-positive rate, the tool clears 99.6% of genuine pre-LLM scientific text while recovering 85% of AI academic text.

The team used papers submitted in 2021 and 2022, before ChatGPT existed, as ground-truth human writing. They set the flag threshold so that exactly 0.4% of those pre-ChatGPT papers would trigger the detector. This created a built-in control: if the rise in flagged papers were simply an artifact of the detector, then 2021 and 2022 papers would show similarly high flag rates as 2026 papers. The data shows they do not.

What Was Measured

The researchers sampled ten field groups, taking roughly 25 papers per field per month from January 2023 to July 2026. They also included eight control months from 2021 and 2022, for a total of 12,750 papers. For each paper, they pulled the version-1 PDF, so that a paper revised in 2026 could not leak modern text back into its 2023 slot.

Importantly, the team scored the full body text rather than just the abstract. They noted that abstracts understate the signal, as they have seen the same paper score under 20% on its abstract and over 70% on its body. Every reported figure carries a bootstrap 95% confidence interval.

Results: Two Waves of Growth

The flagged share of papers remained flat at 0.4% through 2021 and 2022. It lifted off within months of ChatGPT's release and climbed in two waves to about 32% over the most recent complete quarter, peaking near 39% in early 2026.

The spread across fields is large. Computer science leads at about 65%, followed by quantitative biology at 56.3%, electrical engineering and systems at 51.3%, and economics and finance at 47.0%. Applied physics sits at 34.0%, statistics at 31.3%, condensed matter at 24.0%, high-energy physics at 14.0%, and astrophysics at 10.7%. Mathematics is the lowest at 0.7%.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

The control column shows each field's 2021 to 2022 flag rate averaged over three sensitivity settings. The fields that rose the most are not the ones with the highest pre-LLM control level, so an elevated starting point does not explain the rise.

Limitations of the Study

The researchers provided an honest account of the study's limitations. First, each field's pre-ChatGPT control is only 200 papers. At a 0.4% flag rate, only eight papers flag across the entire 2,000-paper control, spread thinly over ten fields. This makes the per-field control rates only approximate. A larger control would not fix this, as pinning a fraction-of-a-percent rate per field would require thousands of control papers per field that pre-2023 arXiv volume does not contain.

Second, a low score can indicate either low adoption of AI or a blind spot in the detector. Mathematics is the clearest case. Mathematics papers are dominated by notation and theorem-proof structure. Once equations and references are removed, the remaining prose is sparse and unlike the scientific English the detector was trained on. A mathematics paper drafted with heavy model assistance may score low because its prose is out of distribution for the detector. So a low score in mathematics is weak evidence that a human wrote the paper. The result is consistent with two very different explanations: lower adoption or reduced detector sensitivity in that register. The data cannot separate them. The fields with the strongest in-distribution assumption, the prose-heavy ones, are also the ones that rise most, so this confound does not account for the aggregate trend. But in the low-scoring fields, the ranking should be read as a lower bound on adoption.

Third, the detector is more sensitive to some generators than others, and the researchers cannot evaluate it against the exact, private mixture of models and prompts that authors actually use. Incomplete coverage lowers the flag rate, so the reported prevalence is a lower bound. The true share is at least what was measured.

Fourth, a flag is not authorship. The detector estimates whether text reads as machine-written, at a calibrated probability with a known error rate. It cannot separate a lightly-edited document from a wholly-generated one, and a single score is never grounds to accuse a specific person. The study reports the prevalence of machine-like writing, which includes heavy AI-assisted editing.

Try the Detector

The detector is cheap to run and the researchers make no money from it. It can be tried for free on any arXiv paper or on custom text.

Related on Neura Market

More from Neura News

AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Jul 21·5 min read
AI Models

Alibaba Qwen-Image-3.0 renders infographics and tiny text in one pass

Alibaba's Qwen team released Qwen-Image-3.0, an image generator designed for practical applications like newspaper layouts and complex infographics. The model processes prompts of up to 4,500 tokens and can render legible text as small as ten pixels, mathematical formulas, and twelve languages in a single pass. It is currently available through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon.

Jul 21·4 min read
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google DeepMind has introduced three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and multimodal performance with 17% fewer output tokens and lower cost. The 3.5 Flash-Lite is the fastest in its series at 350 output tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber, fine-tuned for cybersecurity, will be available exclusively to governments and trusted partners via the CodeMender agent.

Jul 21·6 min read