unslop study flags rising AI-style text on arXiv, with computer science at 65% and math at 0.7%

unslop study flags rising AI-style text on arXiv, with computer science at 65% and math at 0.7%

N
News Editor
2026-07-27 08:15:09
A study by unslop has sparked debate over how much academic writing on arXiv now carries detectable AI-style language. The team examined 12,750 papers across 10 disciplines from January 2023 to July 2026, using pre-ChatGPT papers from 2021 and 2022 as a baseline. It found that the overall flag rate was stable at 0.4% in the control period, then climbed sharply after ChatGPT’s release, reaching 32% in the latest full quarter and nearly 39% at a peak in early 2026. The disciplinary gap was wide. Computer science posted the highest flagged rate at 65%, followed by quantitative biology at 56.3%, electrical engineering at 51.3%, and economics and finance at 47%. Mathematics stood at 0.7%, though the report says that figure is hard to interpret because the detector relies on prose patterns and may struggle with formula-heavy papers. The report also stresses that the tool detects AI-like writing traits, not authorship in a strict sense. It cannot tell whether a paper was lightly polished by an LLM or written extensively by one. The original article also cited Stanford research showing growing LLM use in peer review and paper abstracts, especially in faster-moving and more competitive fields.
arXivAI detectionChatGPTcomputer sciencemathematicsacademic writingStanfordlarge language models

A study from unslop has triggered a fresh argument over AI use in academic writing after reporting that a large share of recent arXiv papers carry detectable machine-style language. In the sample it analyzed, 65% of new computer science papers were flagged by the detector, while mathematics came in at just 0.7%.

The team examined 12,750 arXiv papers spanning January 2023 through July 2026. It focused on 10 disciplines and sampled about 25 full-text papers per field each month. For a control group, the researchers used papers from 2021 and 2022, before ChatGPT was released.

According to the study, the flagged rate during 2021-2022 was stable at 0.4%. After ChatGPT arrived, the curve rose within a matter of months. In the most recent full quarter, the flagged rate reached 32%, and the peak in early 2026 came close to 39%.

unslop study flags rising AI-style text on arXiv, with computer science at 65% and math at 0.7% 3

Large gaps appeared across disciplines

Computer science recorded the highest rate at 65%. The report said the field’s baseline before ChatGPT had been just 0.2%.

Behind computer science were quantitative biology at 56.3%, electrical engineering at 51.3%, economics and finance at 47%, applied physics at 34%, statistics at 31.3%, condensed matter physics at 24%, high-energy physics at 14%, and astrophysics at 10.7%. Mathematics ranked last at 0.7%.

Using the same detector on the same paper repository produced a gap of nearly 100 times between some fields. The study added that these figures should be treated as a lower bound because sufficiently concealed AI-written text may escape detection.

unslop study flags rising AI-style text on arXiv, with computer science at 65% and math at 0.7% 4

Why mathematics may be hard to read for the detector

unslop said the 0.7% result for mathematics does not settle the question. It could mean mathematicians use AI less, or it could mean the detector struggles with the structure of math papers.

A typical mathematics paper is filled with symbols, formulas, theorems, and proofs, leaving relatively little continuous English prose. The detector, by contrast, looks at sentence structure, word choice, and paragraph organization. The limited text that does appear in mathematics papers often follows rigid forms such as setting up assumptions or deriving a result from an earlier lemma, which may not resemble the scientific English seen during the detector’s training.

The researchers listed several limitations. First, the mathematics figure cannot yet separate genuinely lower AI use from detector blindness. Second, the control set was small: each discipline had only 200 pre-ChatGPT papers. At a 0.4% false-positive rate, only 8 out of 2,000 papers would be flagged, which the article described as no more than a rough estimate. Third, the detector cannot catch every instance of AI-written or AI-edited text, so the reported shares are lower-bound figures rather than a full count.

unslop study flags rising AI-style text on arXiv, with computer science at 65% and math at 0.7% 5

Stanford work points to wider LLM use in research workflows

The original article also pointed to a separate two-year line of work from Stanford on who is using AI in academic settings.

In March 2024, a team led by Weixin Liang at Stanford examined peer review and reported that about 10.6% of sentences in ICLR 2024 reviews had been substantially revised by large language models.

In August 2025, the same team expanded the sample to 1.12 million papers, with the study appearing in Nature Human Behaviour. It found that, as of September 2024, the share of computer science abstracts modified by LLMs had reached as high as 22.5%. The article said the heaviest use came from researchers posting preprints more frequently and working in more competitive areas.

unslop study flags rising AI-style text on arXiv, with computer science at 65% and math at 0.7% 6

That pattern, as framed in the article, suggests that faster and more crowded fields are leaning harder on AI tools to save time.

The detector measures AI-like style, not proof of authorship

unslop also cautioned against treating the 65% figure as a direct accusation. The detector cannot distinguish between a paper that had its grammar polished by AI and one that was drafted extensively by a model. What it measures is the presence of machine-like textual traits.

In that sense, saying 65% of papers were flagged is not the same as saying 65% were written by AI from start to finish. It means those papers registered a detectable AI-like style under the tool’s criteria.

unslop study flags rising AI-style text on arXiv, with computer science at 65% and math at 0.7% 7

The article added an example from researchers who ran older papers through detection tools and found that 27% to 74% of the content was marked in red, even though those papers had been written before ChatGPT existed.

One explanation is straightforward. Academic writing today often uses standardized phrasing, highly regular structure, and recurring vocabulary around large language models, benchmark testing, and frontier research. Those traits can overlap with the patterns a detector treats as machine-like. At the same time, LLMs themselves were trained on large volumes of scholarly writing, making the line between “sounds academic” and “sounds AI-generated” harder to draw.

“AI flavor” is becoming a new test applied to text

The article argued that this idea of “AI flavor” in writing is moving from intuition to something people want to quantify. unslop’s detector turns that suspicion into a tool that can generate numbers at scale.

unslop study flags rising AI-style text on arXiv, with computer science at 65% and math at 0.7% 8

Even so, the piece noted that no one has precisely measured what this so-called AI flavor is actually made of. In academic publishing, that uncertainty is changing how text is judged. Writing something carefully used to carry its own credibility. Now, polished wording and orderly structure may invite a different first question: whether AI was involved.

The original piece was published via the WeChat account New Intelligence and credited to ASI Qishilu. MarsBit republished the article.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
300

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.