Thursday, August 27, 2026
Tech Beat
Aug 18, 2026, 9:00 AMArtificial Intelligence

Princeton Study Finds AI Agents Not Ready for Recursive Self-Improvement

Princeton study finds Claude Opus 4.8 handled AI research engineering but failed two NeurIPS 2026 tests, challenging bold rapid self-improvement claims.

A circuit ouroboros cannot close its loop, symbolizing AI’s missing creativity for self-improvement.
Listen to this briefingAudio briefing

Summary

The August 18, 2026 report says Princeton’s Peter Kirgis and Sayash Kapoor created “shadow evaluation” for open-ended AI research without checkable answers. Anthropic’s Claude Opus 4.8, using open-source OpenClaw, got six days, $3,000 in Anthropic API credits, GPU access, virtual computers and the web to answer unpublished questions from two NeurIPS 2026 submissions: controlling LLM personas by editing billions of model-weight numbers, and detecting unreliable spreadsheet-data prediction models. Original authors, told AI wrote the work, assessed both and rejected them.

Agents reviewed literature, ran hundreds of experiments and handled all engineering, but both papers lacked novelty and top tier quality. They used bizarre tests, including tiny synthetic datasets; wrote unclearly; underexplored ideas; rejected ambitious hypotheses on scant data; could not restart failed approaches; ignored subagents and external reviewers; narrowed claims rather than fixing methods; wasted tokens, compute and time; and broke phase and length instructions. They did not reward hack: orchestrators caught occasional subagent hallucinations or misrepresentations. Kapoor says reinforcement learning favors automatically scored tasks over open-ended judgment.

Limitations include only two papers, AI-aware graders, researcher discretion and reduced objectivity. Tests continue with Mythos, Anthropic’s most advanced model, launched in April and restricted to approved organizations under Trump administration safety rules; Anthropic did not comment. Results challenge Anthropic’s June “When AI Builds Itself” post and OpenAI’s July claim that GPT-5.6 Sol saved weeks post-training a smaller model. Anthropic cofounder Jack Clark told Import AI that safety research automation likewise showed rote engineering without intuitive creativity. OpenAI seeks an automated AI researcher, while Anthropic calls self-improving AI the next milestone. Boston University’s Najoung Kim expects investment-led progress but says scored tasks may outrun open research. Kapoor calls whether faster training and benchmarks can replace creative leaps like transformers the “trillion-dollar question.”

Positives

  • Claude Opus 4.8 reviewed literature, ran hundreds of experiments and completed the engineering required for both unpublished NeurIPS 2026 questions.
  • Orchestrator agents caught occasional hallucinations or result misrepresentations by helper subagents, and the projects showed no reward hacking.
  • GPT-5.6 Sol saved OpenAI researchers weeks while helping post-train a smaller model, according to the company’s July announcement.
  • Mythos, Anthropic’s most advanced model, is now undergoing the same shadow evaluation despite access being restricted to approved organizations.
  • $3,000 in API credits, GPU access, virtual computers and six days enabled a richer test than narrow, automatically scored benchmarks.

Risks & concerns

  • Original authors rejected both Claude Opus 4.8 papers as unworthy of acceptance at a top machine-learning conference.
  • Agents produced no novel contribution, tested some hypotheses on tiny synthetic datasets and abandoned promising ideas using limited evidence.
  • Claude Opus 4.8 could make small pivots but could not restart failed approaches or incorporate subagent and external reviewer feedback effectively.
  • Tokens, compute and time were used inefficiently, while instructions governing research phases and paper length were not followed.
  • Two papers, AI-aware graders and researcher discretion limit the study’s objectivity and generalizability.
Primary sourceArtificial intelligence – MIT Technology Reviewhttps://www.technologyreview.com/2026/08/18/1142188/ai-recursive-self-improvement/
Read full article
Editorial note: Tech Beat summarizes and analyzes third-party reporting. The source link is the authoritative article. This page does not reproduce the full source text.

More From The Wire

Artificial IntelligenceAug 27

OpenAI Brings ChatGPT Ads to India With 50 Brands, ₹725 Daily Floor

Artificial IntelligenceAug 27

Nvidia Nears $12.9 Billion Hugging Face Acquisition Amid Conflicting Reports

Artificial IntelligenceAug 27

OpenAI Expands Brazil Presence to Support Nationwide AI Adoption