Transactions on Machine Learning Research (TMLR) says a recent editorial test found that some authors could not explain papers carrying their own names. In a Sept. 16 post on the journal’s official blog, co-editor-in-chief Nihar B. Shah of Carnegie Mellon University (CMU) said that during his Aug. 14-28, 2026 editor rotation, he informally selected 10 submissions that were already on track for desk rejection and invited the authors to discuss them. All 10 were ultimately rejected.

Shah framed the exercise with simple questions: What is the paper’s problem setup? What does this symbol mean? Where, exactly, is the abstract’s claimed result established in the main text? For someone who actually wrote the paper, he said, those questions should be easy to answer without preparation.
Gautam Kamath, also an editor at TMLR, reposted the piece on X and called it a "heroic experiment," saying it confirmed what many had already suspected: some submitters do not really know what is in their own papers.
What happened to the 10 submissions
According to Shah, he contacted authors through OpenReview with a short message saying an editor wanted to speak before sending the paper out for review, in order to better understand the submission, and asked them to email available times.
The responses split quickly. One paper was withdrawn outright. One author said they were too busy to take a call. Meetings were scheduled for the remaining eight, but one author did not show up.
That left seven papers whose authors actually met with Shah. The group was mixed: undergraduates, master’s students, PhD students, faculty members, and independent researchers. Most of the papers were single-author submissions, though not all of them were.
Shah’s questions fell into two broad categories:

- basic questions about the problem setting, notation, and the results claimed in the paper;
- detail questions about technical expressions, theoretical results, and why specific experimental design choices were made.
Only one of the seven authors answered everything. Three more could describe the high-level idea, but ran into trouble once Shah pressed on technical specifics. The remaining three performed worst: they could not answer even basic questions, and all three papers were single-author papers.
Shah wrote that two of those authors appeared to have almost no substantive understanding of their own manuscripts. A third could not identify where several key results claimed in the abstract were actually presented or supported in the body of the paper.
Even the strongest case still failed. Shah said he found a major error in one of that paper’s main conclusions, and the author later acknowledged it. TMLR desk-rejected the submission, but said the author could resubmit after correcting the error or narrowing the conclusion. The other nine papers were desk-rejected with no resubmission path.
Two follow-up episodes after the meetings
Shah also described two episodes after the interviews.
In the first, authors of two papers could not answer basic questions during their meetings but later sent written replies by email. Shah ran both emails through the AI text detection tool Pangram. Both came back as "100% AI."
In the second, one author tried to explain an analysis method used in the paper but ended up, in Shah’s telling, describing a full p-hacking workflow — repeatedly adjusting the analysis until the data turned out "significant."

Shah added that he is not opposed to AI itself. To get through eight papers in two weeks, he said, he also used large language models to help him understand submissions and even learn concepts he did not already know. The question for him is not whether AI was used. It is whether authors can take responsibility for work published under their names.
Why TMLR is doing this
TMLR launched in 2022 and is run by the team behind the Journal of Machine Learning Research (JMLR). The journal is known for a review standard that focuses on whether conclusions are supported by evidence, rather than treating novelty or state-of-the-art performance as grounds for rejection. Its reviewers, action editors, and editors-in-chief all serve as unpaid volunteers, which has made the recent submission wave harder to manage.
In a June notice, TMLR said submissions over the prior year had tripled, while single-author submissions had increased 13-fold. The editorial team had even seen cases where one person submitted five papers in a single day. Shah gave another number in the new post: TMLR’s desk-rejection rate was about 6% in 2023, but has now risen to about 53%, meaning more than half of submissions do not make it to external review.
TMLR has rolled out several measures this summer in response.
Annual author submission quotas
The first is an annual submission quota tied to authors. Instead of using a flat cap of "at most N papers per person," TMLR adopted what it described as a harmonic-style quota rule, under which the quota consumed by each author on a paper declines as the number of co-authors rises.
- Authors who submit only single-author papers can submit at most two papers per year.
- If all submissions are nine-author papers, the yearly limit is nine papers.
- Active reviewers and action editors get double quota.
The rule took effect on July 1. In the journal’s explanation, it did not simply divide quota by headcount because that would create an incentive to add nominal co-authors just to gain more submission capacity.

AI-generated reviews for reliability
The second is AI review. In July, TMLR said every submission would receive, alongside normal human review, an AI-generated review focused only on "reliability" — whether the paper’s conclusions are backed by accurate, clear, and convincing evidence. The AI review does not make subjective judgments and does not recommend acceptance or rejection. The final decision stays with the action editor. After evaluation, TMLR chose the AI reviewer built by CSPaper.
Clear writing added to acceptance criteria
The third is a stronger writing requirement. On Aug. 28, TMLR revised its existing "interest to the audience" criterion and explicitly required papers to communicate their findings clearly to readers. The announcement said current AI writing often wanders, piles on terminology, and is hard to follow. If a paper is written almost entirely by AI with little human involvement, TMLR said, it is unlikely to meet that standard at the current state of these systems.
Shah wrote that the author interviews gave the editorial team more confidence in its desk-rejection process and suggested that the quota policy and the emphasis on clear writing were helping.
TMLR is not alone
The problem, Shah argued, is not unique to TMLR. Other major AI publication venues have been dealing with similar issues this year.
In early June, the official NeurIPS blog said that of 971 submissions to the NeurIPS 2026 Position Paper track, 273 — or 28.2% — were scored by Pangram as 100% AI. That track explicitly required papers to be "substantially written by humans," with AI limited to peripheral editing tasks such as language polishing.
NeurIPS also said the rise in AI-written submissions was broader than one track. In the Datasets and Benchmarks track, the number of papers with Pangram scores of at least 90% increased more than tenfold from 2025 to 2026. At the same time, NeurIPS acknowledged that detection results are sensitive to parameters: when a medium-size detection window was used instead, the share of papers scoring between 90% and 100% AI dropped from 42.7% to 12.7%.

Earlier, in mid-May, Thomas G. Dietterich, who oversees the computer science section at arXiv, said on social media that if there is "clear evidence" that authors failed to check large-model output in a submission — for example, hallucinated references or chatbot dialogue left in the manuscript — all authors would be barred from submitting to arXiv for one year. After that, future submissions would have to be accepted by a formal peer-reviewed venue before appearing on arXiv. Dietterich said authorship means every listed author is responsible for the entire paper, no matter how the content was produced.
Comparable signals have appeared outside computer science. A Columbia University study published in The Lancet found that among biomedical papers indexed by PubMed Central, the share containing at least one fabricated citation rose from about 4 in 10,000 in 2023 to about 57 in 10,000 by early 2026.
From checking text to checking people
Shah acknowledged that his experiment took 20 to 25 hours and still covered only eight papers, making it difficult to scale under current submission volumes.
He said his team is looking at alternatives that can scale. Just days earlier, Shah and CMU researchers Justin Payan, Bálint Gyevnár, and Atoosa Kasirzadeh released a paper proposing an evaluation method called greCAPTCHA. The name combines GRE, the graduate admissions exam, with CAPTCHA, the test used to distinguish humans from machines.
The idea is not to keep testing whether a text was written by AI, but to examine the author directly. An author submits the paper together with a personal contribution statement. The system then generates questions automatically, the author answers under proctored conditions, and the process produces an evaluation report that could be used by journals, employers, or admissions offices. The capability being tested is what the researchers call the "capacity to verify" — whether an author has the knowledge and reasoning ability needed to critically evaluate the content they contributed.
The questions fall into four categories:

- identifying the real data reported in the paper from among multiple versions, or seeded-error recognition;
- explaining design choices whose reasoning is not spelled out in the paper;
- explaining background concepts that the paper assumes readers already know;
- identifying the specific conditions under which the method would fail.
The team tested 31 researchers, with each participant answering questions about one of their own papers and one unfamiliar paper. The system reached an AUC of 0.90 in distinguishing "own paper" from "unfamiliar paper," and the AUC rose to 0.934 when multiple-choice questions were removed. The multiple-choice portion alone had little discriminative value, with an AUC of 0.595, because the answers could often be searched directly in the PDF.
Another result stood out: there was no statistically significant difference between scores on unfamiliar papers from the participant’s own field and unfamiliar papers from outside the field. That suggests the test is not measuring field knowledge alone.
One example in the paper captured the issue neatly. A participant fed the unfamiliar paper to an AI system before the test and had it explain the manuscript from start to finish, yet still could not answer deeper questions. That participant scored 1.3 on the unfamiliar paper and 51 on their own paper.
The method is still far from finished. The paper reports a lowest error rate of about 20%, meaning roughly one in five genuine authors could be misclassified. Participants complained most often about scoring: some said it was unclear how detailed an answer needed to be, while others said the rubric was too rigid. One author, reacting to a scoring rule that rejected their response, said: "I designed that thing myself."
Other participants noted that in collaborative papers, some sections were written by co-authors and they would not be able to answer questions about those parts. There were also concerns that this kind of test could end up measuring English fluency or typing speed, making it unfair to non-native English speakers and disabled researchers.
What Shah says this means
Shah closed with several takeaways, including two he highlighted in particular.

The first concerns credit. In today’s research system, academic contribution is recognized mainly through authorship. If a listed author cannot explain or defend the paper, then publication becomes a much weaker signal of that person’s contribution as a researcher.
The second concerns the peer-review cycle itself. Many journals and conferences ask authors who submit papers to review other people’s work. If authors do not understand their own manuscripts, their ability to review others’ papers naturally comes into question. With AI review and AI writing entering the process at the same time, Shah warned that the foundations of peer review are at risk if that cycle breaks down.
He also stressed the limits of the experiment. The sample size was just 10 papers, selected informally, and all came from submissions that were already headed for desk rejection. The exercise shows what can be found among papers already judged likely to be low quality. It does not describe the overall composition of TMLR submissions. Shah said as much himself: one purpose of the experiment was to test whether the journal’s desk-rejection process was working.
Still, the line drawn by the experiment was simple. A paper may be produced with help from AI, but the people whose names are on it must be able to explain what it says.
This article is based on a report by the WeChat public account "机器之心" (ID: almosthuman2014), authored by 机器之心.

