Researchers have analyzed the full text of 12,750 arXiv papers and concluded that approximately one third of new submissions appear to be machine-written. The study, which covers papers from January 2023 to July 2026, provides a detailed look at how AI writing has spread across academic fields since the release of ChatGPT.
Methodology and False-Positive Floor
The researchers built their study around a common criticism of AI detectors: that they often flag genuine human writing as machine-generated. To address this, they calibrated their detector specifically for academic writing. At a 0.4% false-positive rate, the tool clears 99.6% of genuine pre-LLM scientific text while recovering 85% of AI academic text.
The team used papers submitted in 2021 and 2022, before ChatGPT existed, as ground-truth human writing. They set the flag threshold so that exactly 0.4% of those pre-ChatGPT papers would trigger the detector. This created a built-in control: if the rise in flagged papers were simply an artifact of the detector, then 2021 and 2022 papers would show similarly high flag rates as 2026 papers. The data shows they do not.
What Was Measured
The researchers sampled ten field groups, taking roughly 25 papers per field per month from January 2023 to July 2026. They also included eight control months from 2021 and 2022, for a total of 12,750 papers. For each paper, they pulled the version-1 PDF, so that a paper revised in 2026 could not leak modern text back into its 2023 slot.
Importantly, the team scored the full body text rather than just the abstract. They noted that abstracts understate the signal, as they have seen the same paper score under 20% on its abstract and over 70% on its body. Every reported figure carries a bootstrap 95% confidence interval.
Results: Two Waves of Growth
The flagged share of papers remained flat at 0.4% through 2021 and 2022. It lifted off within months of ChatGPT's release and climbed in two waves to about 32% over the most recent complete quarter, peaking near 39% in early 2026.
The spread across fields is large. Computer science leads at about 65%, followed by quantitative biology at 56.3%, electrical engineering and systems at 51.3%, and economics and finance at 47.0%. Applied physics sits at 34.0%, statistics at 31.3%, condensed matter at 24.0%, high-energy physics at 14.0%, and astrophysics at 10.7%. Mathematics is the lowest at 0.7%.
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
The control column shows each field's 2021 to 2022 flag rate averaged over three sensitivity settings. The fields that rose the most are not the ones with the highest pre-LLM control level, so an elevated starting point does not explain the rise.
Limitations of the Study
The researchers provided an honest account of the study's limitations. First, each field's pre-ChatGPT control is only 200 papers. At a 0.4% flag rate, only eight papers flag across the entire 2,000-paper control, spread thinly over ten fields. This makes the per-field control rates only approximate. A larger control would not fix this, as pinning a fraction-of-a-percent rate per field would require thousands of control papers per field that pre-2023 arXiv volume does not contain.
Second, a low score can indicate either low adoption of AI or a blind spot in the detector. Mathematics is the clearest case. Mathematics papers are dominated by notation and theorem-proof structure. Once equations and references are removed, the remaining prose is sparse and unlike the scientific English the detector was trained on. A mathematics paper drafted with heavy model assistance may score low because its prose is out of distribution for the detector. So a low score in mathematics is weak evidence that a human wrote the paper. The result is consistent with two very different explanations: lower adoption or reduced detector sensitivity in that register. The data cannot separate them. The fields with the strongest in-distribution assumption, the prose-heavy ones, are also the ones that rise most, so this confound does not account for the aggregate trend. But in the low-scoring fields, the ranking should be read as a lower bound on adoption.
Third, the detector is more sensitive to some generators than others, and the researchers cannot evaluate it against the exact, private mixture of models and prompts that authors actually use. Incomplete coverage lowers the flag rate, so the reported prevalence is a lower bound. The true share is at least what was measured.
Fourth, a flag is not authorship. The detector estimates whether text reads as machine-written, at a calibrated probability with a known error rate. It cannot separate a lightly-edited document from a wholly-generated one, and a single score is never grounds to accuse a specific person. The study reports the prevalence of machine-like writing, which includes heavy AI-assisted editing.
Try the Detector
The detector is cheap to run and the researchers make no money from it. It can be tried for free on any arXiv paper or on custom text.

