Thursday, August 27, 2026
Tech Beat
Aug 26, 2026, 9:00 AMArtificial intelligence

AI Puzzle Tests Reveal Rapid Gains and Persistent Reasoning Flaws

AI puzzle tests expose sharp gains and stubborn flaws in spatial, visual and complex reasoning, despite near-perfect Connections results in early 2025.

Listen to this briefingAudio briefing

Summary

Puzzles have benchmarked AI since IBM scientist Arthur Samuel popularized “machine learning” in a 1959 paper on a checkers algorithm. Columbia University scientists found leading models solved only 18% of New York Times Connections puzzles in late 2024, yet some approached perfect performance by early 2025, showing unusually rapid gains.

Persistent failures expose a gap between recall and adaptable reasoning. Vision language models struggle with mental rotation and manipulating 3D objects; Google and University of Illinois Urbana Champaign researchers found in 2024 that slight changes to Knights and Knaves puzzles trigger memorized answers. SimpleBench likewise catches frontier models on easy questions with trick phrasing. ARC AGI performance improves when colored grids are encoded as number strings, and correct answers can rely on complex rules that do not generalize, unlike humans’ simpler visual concepts. Humans remain more vulnerable to intuitive math and wording traps, while models answer more deliberately. Apple researchers found LLMs solve simple Tower of Hanoi and river crossing tasks but falter at six or more disks or people. University of Washington, Stanford University and Allen Institute for AI researchers found similar scaling limits in logic grids, although critics dispute whether rising errors indicate an LLM specific defect or ordinary difficulty under complexity.

Positives

  • By early 2025, some models solved New York Times Connections puzzles almost perfectly, up from an 18% success rate for leading models in late 2024.
  • ARC AGI performance has improved substantially, despite the benchmark’s demands for abstract and visual rule inference.
  • LLMs often answer intuitive math and wording traps more deliberately than humans, avoiding some common cognitive biases.
  • Apple researchers found models can master simpler Tower of Hanoi and river crossing problems.

Risks & concerns

  • Vision language models still perform extremely poorly on mental rotation and other tasks requiring manipulation of 3D objects.
  • A 2024 Google and University of Illinois Urbana Champaign study found small Knights and Knaves variations can provoke memorized rather than adapted answers.
  • SimpleBench exposes frontier models failing easy questions whose wording resembles harder training examples.
  • Correct ARC AGI answers may rely on complicated rules that do not generalize, rather than the simple visual concepts humans use.
  • Tower of Hanoi and river crossing performance deteriorates when problems reach six or more disks or people.
Primary sourceArtificial intelligence – MIT Technology Reviewhttps://www.technologyreview.com/2026/08/26/1141952/puzzles-ai-models-flub-these-tests/
Read full article
Editorial note: Tech Beat summarizes and analyzes third-party reporting. The source link is the authoritative article. This page does not reproduce the full source text.

More From The Wire

Artificial intelligenceAug 20

Virgin Atlantic Uses Generative AI Market Models for Dynamic Pricing

A blue toy robot’s heart cord connects to a crumbling cloud, symbolizing the fragility of AI companionship. Artificial intelligenceAug 17

Moxie Shutdowns Expose Risks of AI Robot Friends for Autistic Children

A silicon seed grows into an Indonesia-shaped canopy protecting lungs, rice and seismic waves. Artificial intelligenceAug 14

Indonesia Opens First University AI Center With UGM, Indosat and NVIDIA