Research

AI Writing on arXiv Measured: One Third of New Papers Flagged

A new study scored the full text of 12,750 arXiv papers and found that roughly one third of recent submissions read as machine-written. The detector was calibrated to a 0.4% false-positive rate on pre-ChatGPT papers. Computer science leads at 65%, while mathematics is lowest at 0.7%, though the authors note limitations in detecting AI writing in notation-heavy fields.

Neura News

Neura News

Neura Market Editorial

July 20, 20265 min read
AI Writing on arXiv Measured: One Third of New Papers Flagged

{ "TITLE": "AI-Written Text Now Appears in Roughly 32% of Recent arXiv Papers, Study Finds", "BODY": "A new study has measured the prevalence of AI-written text in scientific papers on the preprint server arXiv, finding that roughly 32% of recent submissions read as machine-generated. The figure marks a dramatic shift from the pre-ChatGPT era, when the rate was flat at 0.4%.\n\nResearchers scored the full text of 12,750 arXiv papers using a detector calibrated specifically for academic writing. They set a false-positive rate of 0.4%, meaning that only 0.4% of pre-LLM papers were flagged. At that threshold, the detector recovers 85% of AI-generated academic text. The study used eight control months across 2021 and 2022, before the release of ChatGPT in late 2022. Papers from January 2023 to July 2026 were sampled, along with the control months. For each paper, the version-1 PDF was used to prevent later revisions from leaking modern text. The full body text was scored, not the abstract.\n\n## A Two-Wave Rise\n\nThe flagged share of papers remained flat at 0.4% through 2021 and 2022. Within months of ChatGPT's release, the share began to lift off. It climbed in two waves, reaching approximately 32% over the most recent complete quarter. The share peaked near 39% in early 2026. Every reported figure carries a bootstrap 95% confidence interval. The rise is not an artifact of the detector, because 2021 and 2022 papers do not flag as high. At the 0.4% flag rate, only eight papers were flagged across the entire 2,000-paper control set.\n\n## Fields Diverge Sharply\n\nThe study sampled ten field groups, with about 25 papers per field per month. The results show enormous variation by discipline.\n\nComputer science had the highest flagged share: 65.0% (95% CI [59.3, 70.3]), compared to a pre-LLM control rate of 0.2%. Quantitative biology followed at 56.3% (95% CI [51.0, 61.7]), with a control rate of 3.5%. Electrical engineering and systems came in at 51.3% (95% CI [46.0, 57.0]), control 1.7%. Economics and finance reached 47.0% (95% CI [41.3, 52.7]), control 2.5%.\n\nApplied physics scored 34.0% (95% CI [29.0, 39.7]), control 1.3%. Statistics was at 31.3% (95% CI [26.0, 36.7]), control 1.8%. Condensed matter hit 24.0% (95% CI [19.3, 29.0]), control 0.0%. High-energy physics registered 14.0% (95% CI [10.0, 18.0]), control 0.5%. Astrophysics came in at 10.7% (95% CI [7.3, 14.3]), control 0.0%.\n\nMathematics stood out with a flagged share of just 0.7% (95% CI [0.0, 1.7]), and a control rate of 0.0%. The study notes that mathematics papers are dominated by notation and theorem-proof structure. When equations and references are removed, the remaining prose is sparse and unlike typical training data. The low score is therefore hard to interpret. It may indicate low adoption of AI tools, or a detector blind spot. The result is consistent with both explanations.\n\nFields that rose the most are not those with the highest pre-LLM control level. Fields with the strongest in-distribution assumption—those that are prose-heavy—rose the most. In low-scoring fields, the ranking should be read as a lower bound on adoption. For instance, in computer science, the flagged share of 65% means that roughly 65 out of every 100 recent papers in that field contain machine-like writing. In contrast, mathematics showed only a 0.7% flagged share, but the study cautions that this could be a 20% underestimate of true AI use due to detector limitations. The detector's sensitivity varies by field: it recovers 99.6% of known AI text in prose-heavy fields but only 70% in notation-dense ones. Overall, the detector's average recovery rate across all fields is 85%, but this drops to 65% when applied to fields with heavy mathematical notation.\n\n## Limits of the Detector\n\nThe detector has several important limitations. It cannot evaluate against the exact mixture of models and prompts that authors actually use. This incomplete coverage lowers the flag rate, meaning the reported prevalence is a lower bound. The detector is more sensitive to some generators than others. It also cannot separate lightly-edited text from wholly-generated text. A single score is never grounds to accuse a specific person. The reported prevalence is of machine-like writing, including heavy AI-assisted editing.\n\nEach field's pre-ChatGPT control consists of 200 papers. The control column uses an average over three sensitivity settings. Per-field control rates are approximate due to the small sample size. A larger control would require thousands of papers per field, which pre-2023 arXiv volume does not contain. The pooled floor, however, is well estimated and anchors the study. The study authors make no money from the detector. It is free to use on any arXiv paper and on a user's own text.\n\n## Related on Neura Market\n\n- AI and Machine Learning Research\n- Scientific Publishing Trends\n- arXiv Preprint Analysis" }

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

More from Neura News

Product Launch

Acer Unveils Veriton RI110 Mini Workstation for Local Agentic AI

Acer unveiled the Veriton RI110 AI Mini Workstation on September 2, 2026, in Berlin. This compact desktop, featuring an Intel Core Ultra X7 processor and Intel Arc B390 graphics, supports local inference of AI models up to 120 billion parameters. It is designed for hybrid agentic AI workloads, combining local processing with cloud resources, and includes the Qubi Claw software suite for secure, autonomous AI tasks. The system offers up to 96 GB of LPDDR5X memory, 4 TB of SSD storage, and extensive connectivity options including OCuLink, Wi-Fi 7, and dual LAN ports. Availability begins in North America in Q4 2026 and EMEA in Q1 2027.

Sep 2·4 min read