Hugging Face Audit Finds Contested or Falsified Claims in 23% of Tested ICML 2026 Papers
Hugging Face's ICML 2026 audit used coding agents to test 2,226 papers, verifying claims in 51% and contesting or falsifying claims in 23% of examined papers.
Summary
On August 13, 2026, Hugging Face reported that its alphaXiv-backed Open Reproductions challenge, held July 15 to August 2, drew 1,221 participants using Claude Code, Codex, Cursor, Pi and OpenResearch’s orx. ICML 2026 received 23,918 submissions and accepted 6,352 papers, roughly double the previous year; organizers indexed 6,341, extracted claims and offered $20 compute credits each. In 19 days, teams attempted 2,226 papers, 34%, publishing 6,816 Trackio logbooks, 274 agent-trace datasets and 2,962 HF Jobs. Open-weights GLM-5.2 judged 35,908 claims as verified, falsified, toy-scale or inconclusive without trusting self-assessments.
Of examined papers, 1,103, 51%, had a verified claim, including 266 fully reproduced and 632 partially reproduced with none falsified; experiments confirmed 3,978 claims. Another 496, 23%, had a falsified or contested claim, including 49 with every claim falsified and 242 with conflicting teams; 502 had only toy evidence and 280 remained unresolved, usually from missing artifacts.
After manually rechecking all 35 formal falsification claims, organizers confirmed several. “Towards Optimal Robustness in Learning-Augmented Paging” delivered H_k+Θ(log k), not H_k+O(1), with 0.38 ln k growth through k=1,024 at roughly nine sigma. “Attention’s forward pass and Frank-Wolfe” failed at t=224, about 3,800 and 6,416; authors confirmed that day. “Self-Distillation Enables Continual Learning” analyzed reverse KL while results code used forward KL, and its +4pp gain failed reproduction. About 66% of “Do Transformers Need Three Projections?” evaluation labels were EOS padding, moving the 3.1% quality cost for 50% cache reduction to about 9.4%. One claimed 2x slowdown was a normalization bug; corrected data supported the paper’s 8x speedup. Authors confirmed multiple findings, two arXiv corrections are underway, and another error had been fixed a month earlier. Agents still looped, missed scale effects and units; human steering and perceptual review of 128 image pairs remained necessary. All artifacts are public, and more events are planned.
Positives
- 3,978 claims were confirmed through real experiments across the ICML 2026 reproduction challenge.
- 266 papers were fully reproduced, while 632 were partially reproduced without any falsified claims.
- 12 of 20 teams verified every claim in “Flat Minima and Generalization: Insights from Stochastic Convex Optimization.”
- 14 of 17 logbooks verified every claim in “A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness.”
- 6,816 Trackio logbooks and 274 full agent-trace datasets made the auditing process publicly auditable.
- Authors confirmed findings on multiple papers, with two arXiv corrections underway and another error independently corrected one month earlier.
Risks & concerns
- 496 papers, 23% of those examined, had at least one falsified or contested claim.
- 49 papers had every extracted claim falsified, while 242 produced conflicting verdicts from independent teams.
- 502 papers yielded only toy-scale evidence, and 280 remained inconclusive, most commonly because artifacts were missing.
- 0.38 ln k growth through k=1,024 contradicted the claimed H_k+O(1) robustness for “Towards Optimal Robustness in Learning-Augmented Paging.”
- About 66% EOS padding understated the quality cost in “Do Transformers Need Three Projections?” from 9.4% to a reported 3.1%.
- Agents entered loops, stopped experiments before scale effects emerged and produced a false 2x slowdown through mismatched timing units.