ICML 2026's 24,000-Paper Trial Finds LLM Review Policies Barely Move Outcomes, but a Third of Reviewers Broke Them

A randomized experiment across a major machine-learning conference found that banning or permitting limited LLM use made almost no difference to paper scores or decisions, even as many reviewers ignored whichever rule they were given.

EduFabTech · 27 September 2026 · 4 min read · 4 views
A balance scale shows the two ICML review policies landing on nearly identical scores and verdicts, while a cracked lock marks the 36.5% of reviewers who broke their assigned rule.
EduFabTech · Own work

A randomized experiment run inside one of the largest machine-learning conferences in the world has produced a rare thing in the debate over AI and peer review: a controlled answer instead of an anecdote. Researchers working with the organizers of the 2026 International Conference on Machine Learning (ICML) randomly assigned reviewers and papers to either a strict ban on large language model (LLM) use or a policy that allowed limited assistance, then measured what changed. The paper, posted to arXiv on September 16, 2026, reports that almost nothing did — at least not in the outcomes that matter most to authors.

The scale is what makes the result credible rather than anecdotal. ICML 2026 processed 24,661 submissions and 17,886 reviewers, according to the study by Sunnie S. Y. Kim, Wesley Hanwen Deng, Jennifer Wortman Vaughan and colleagues at Microsoft Research, the University of Pennsylvania, Google Research, the University of Wisconsin-Madison, EPFL and Carnegie Mellon University. Within that pool, 8,387 reviewers who said they would accept either policy were randomly split: 1,008 were assigned the conservative, no-LLM-at-all policy, and 912 were assigned the permissive policy that allowed reviewers to use an LLM to help understand a paper or polish their own prose, but not to judge it or draft the review itself.

That distinction between "understand and polish" versus "judge and draft" is exactly what ICML's official policy for LLM use in reviewing lays out. Reviewers under the conservative policy could use non-traditional LLM tools only by accident; those under the permissive policy could not ask an LLM to assess a paper's strengths, weaknesses or significance, or to write review text on their behalf. Program chairs randomized a subset of assignments specifically so the difference between the two policies could be tested rather than assumed.

Three side-by-side bar comparisons show paper rejection rates staying flat between policies while self-reported rule-breaking and Pangram's AI-detection scores diverge sharply.
Three side-by-side bar comparisons show paper rejection rates staying flat between policies while self-reported rule-breaking and Pangram's AI-detection scores diverge sharply.EduFabTech · Own work

Scores, decisions and confidence barely moved

Among reviewers randomized between the two policies, the paper scores they gave differed by 0.002 points on a shared scale, with an effect size (Cohen's dz) of 0.00 and a p-value of 0.96 — statistically indistinguishable from no effect. Reviewer confidence differed by 0.004 points (dz = 0.00, p = 0.92). A parallel comparison at the paper level, where some papers were switched from the permissive to the conservative pool, found final decisions barely differed either: 73.0% of papers were rejected under the switched conservative condition versus 73.5% under the permissive one (p = 0.56). The one measurable difference was length: reviews written under the permissive policy ran about 5.5–7% longer than those under the conservative one, a small but statistically significant gap (p < 0.001 in the paper-level comparison).

Compliance, not policy design, was the weak point

The more striking finding sits in the anonymous post-survey of 1,486 reviewers and area chairs. Despite being told LLM use was strictly prohibited, 22.5% of conservative-policy reviewers admitted using one anyway. Among reviewers given the more permissive policy, 36.5% reported at least one use that the policy explicitly disallowed, such as asking an LLM to identify a paper's weaknesses. Across the full reviewer survey, 50.5% of the 1,311 responding reviewers said they had used an LLM at some point in the review process, and nearly 47% of area chairs said the same.

The authors added an independent check by running submitted reviews through the AI-text detector Pangram. The pattern lines up with the self-reported numbers: reviews from reviewers who kept the conservative policy throughout were classified as 59.0% human-written, compared with 29.0% for reviewers who stayed on the permissive policy for the entire process. In other words, even under a flat ban, a large share of review text carried the statistical signature of LLM involvement.

Two side-by-side scorecards contrast the no-LLM and limited-LLM policies on reviewer counts, rejection rates, rule-breaking admissions, and AI-detected review authorship.
Two side-by-side scorecards contrast the no-LLM and limited-LLM policies on reviewer counts, rejection rates, rule-breaking admissions, and AI-detected review authorship.EduFabTech · Own work

Why the gap between rule and behavior matters

ICML's organizers were already tracking this problem before the randomized results came in. A March 2026 post from the program chairs, scientific integrity chair and communications chairs described watermarking submission PDFs with hidden instructions that only an LLM would follow, which caught 506 reviewers under the conservative policy feeding papers into a chatbot; the conference desk-rejected the papers of 398 of the reviewers involved in reciprocal review pools as a result. That enforcement effort and the new randomized study describe the same underlying pattern from two different angles — one measuring who got caught, the other measuring how much it changed the review itself.

The finding also fits a wider trend outside this one conference. A Frontiers survey of roughly 1,600 academics across 111 countries, reported by Nature in December 2025, found more than half had used AI tools while peer reviewing, often in ways that ran against publisher guidance. Read together with the ICML trial, the picture that emerges is not that LLM policies fail to work because they are poorly designed, but that stated policy and actual reviewer behavior have become two different things — and that, so far, the gap between them hasn't visibly changed which papers get accepted.


References
  1. Sunnie S. Y. Kim, Wesley Hanwen Deng, Jennifer Wortman Vaughan, Buxin Su, Weijie Su, Alekh Agarwal, Sharon Li, Martin Jaggi, Daniel G. Goldstein, Nihar B. Shah, Miroslav Dudík. Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026. arXiv preprint, 2026. doi:10.48550/arXiv.2609.19420
  2. ICML 2026 Program Committee. Policy for LLM Use in Reviewing. International Conference on Machine Learning (ICML), 2026. link
  3. Alekh Agarwal, Miroslav Dudik, Sharon Li, Martin Jaggi, Nihar B. Shah, Katherine Gorman, Gautam Kamath. On Violations of LLM Review Policies. ICML Blog, 2026. link
  4. Nature News (Springer Nature). More Than Half of Researchers Now Use AI for Peer Review — Often Against Guidance. Nature, 2025. doi:10.1038/d41586-025-04066-5