Over 30% of Recent arXiv Submissions Detected as AI-Written, Study Finds
Key Takeaways
- ▸AI-written content in arXiv submissions surged from 0.4% in 2021-2022 to ~32% in recent quarters, with sharp acceleration following ChatGPT's release in late 2022
- ▸Detection rates vary dramatically by discipline—computer science at 65% versus mathematics at 0.7%—suggesting uneven adoption or field-specific detector limitations
- ▸The study's rigorous methodology uses pre-ChatGPT papers as calibrated baselines to rule out false positives, confirming the rise is genuine rather than a measurement artifact
Summary
A comprehensive analysis of over 12,750 arXiv papers has found that approximately 32% of recent submissions now read as machine-written, with detection rates peaking near 39% in early 2026. The study employs a detector calibrated to academic writing with a rigorous 0.4% false-positive rate on pre-ChatGPT papers, tracking submissions across ten field groups from 2021 through July 2026. The research reveals dramatic variation by discipline, with computer science leading at 65% AI-written content and mathematics at just 0.7%, suggesting significant differences in AI adoption or detector sensitivity across fields.
The study's methodology uses papers from 2021-2022 as ground-truth baselines, setting detector thresholds so that only 0.4% of pre-LLM academic writing triggers the AI-written flag. This approach directly addresses the most common criticism of such studies by eliminating false-positive inflation. The research shows a clear inflection point within months of ChatGPT's release, with AI-detected content climbing in two waves over the following three years.
While the findings provide one of the first rigorous measurements of AI adoption in academic writing, the study carefully documents important limitations. Mathematics papers may score low due to detector blind spots caused by heavy notation and sparse prose rather than genuine low adoption. The authors acknowledge that per-field control sample sizes are approximate, and the dramatic variation across disciplines raises open questions about whether differences reflect genuine adoption patterns or measurement artifacts.
- Important limitations remain, including potential detector blind spots in notation-heavy fields and uncertainty about whether low rates indicate low adoption or reduced detector sensitivity
Editorial Opinion
This study provides one of the first empirically rigorous measurements of AI adoption in academic writing, revealing a striking picture of rapid integration—especially in computer science. While the methodology is sound, the dramatic variation across disciplines raises important questions about whether AI adoption is genuinely uneven or whether detection tools themselves have field-dependent blind spots. Academic institutions and journal editors must now urgently develop standards for AI-assisted writing in peer-reviewed research, balancing innovation with scholarly integrity.



