Wednesday, October 7, 2026
Tech Beat
Oct 6, 2026, 4:19 PMArtificial Intelligence

Why AI Fails in Long Conversations and Complex Workflows

Microsoft researcher Jennifer Neville finds AI performance drops in multiturn chats and long workflows, urging user checks while teams develop new fixes.

Listen to this briefingAudio briefing

Summary

On October 6, 2026, Microsoft Principal Applied Scientist Chad Atalla interviewed Jennifer Neville, a Microsoft partner research manager since 2021 and Purdue’s Samuel D. Conte chair professor of computer science and statistics. Across 20 years of teaching and advising, Neville has published more than 130 papers with over 10,000 citations, won an NSF CAREER Award and best paper awards at ICDM and ICLR, and made IEEE’s “10 to Watch” in AI list. Her AI Interaction and Learning team evaluates models in realistic multiturn, collaborative and long horizon work, arguing that simple benchmarks miss failures that matter to users.

By converting public single turn benchmarks into simulated conversations where users add requirements gradually, the team found performance across current models falls significantly versus fully specified prompts, though no percentage was disclosed. It is developing reinforcement learning fixes. Meanwhile, users can restart a confused chat and submit one complete prompt. In repeated document editing and other agentic workflows, errors accumulate, models lose semantic content and struggle to track intent, state, relevance and changes. Open problems include tool use, human oversight, self monitoring and rollback. Microsoft analyzes consumer logs only through restricted, privacy preserving processing, not direct inspection or fine tuning on users’ data. Donated descriptions and detailed feedback can still be distilled into model improvements.

Neville says today’s systems can improve productivity but are not reliable enough for full delegation, so users should verify outputs and retry incorrect answers. Research is testing whether wrappers and reasoning can overcome transformers’ known computational limits or alternative architectures are needed, with longer term world understanding next. She recalls experts at AT&T Labs in 2000 predicting Go would take 50 to 100 years, a solved problem far sooner, suggesting present barriers may also fall quickly.

Positives

  • More than 130 papers and 10,000 citations underpin Jennifer Neville’s research into practical AI behavior and structured data.
  • Realistic evaluations now test multiturn conversations, collaborative environments and long horizon workflows that simple benchmarks overlook.
  • Reinforcement learning methods are being developed to help models follow requirements added across multiple conversational turns.
  • Restricted, privacy preserving systems extract broad patterns from Microsoft consumer logs without researchers directly inspecting the underlying data.
  • Detailed descriptions accompanying negative feedback can be distilled into improvements for later model updates.
  • Restarting a confused chat with one complete prompt can improve results using current models.

Risks & concerns

  • Current models perform significantly worse when one fully specified prompt is divided across multiple turns, though no percentage was disclosed.
  • Uncaught errors accumulate during repeated document edits, increasing confusion and eroding semantic content over long workflows.
  • AI agents still struggle to track user intent, document state, relevant information and changes across complex tasks.
  • Current systems are not reliable enough for full delegation and require human verification of important outputs.
  • Effective tool use, human intervention, self monitoring and rollback remain open problems for agentic systems.
  • It remains unclear whether wrappers can overcome transformer limitations or fundamentally different architectures will be required.
Primary sourceMicrosoft Researchhttps://www.microsoft.com/en-us/research/podcast/what-ai-gets-wrong-and-what-failure-teaches-us/
Read full article
Editorial note: Tech Beat summarizes and analyzes third-party reporting. The source link is the authoritative article. This page does not reproduce the full source text.

More From The Wire

Artificial IntelligenceOct 6

OpenAI Makes ChatGPT textGrain Watermarks Default in EU

Artificial IntelligenceOct 6

Underdog Launches Private On-Device AI With Stripe Fee Model

Artificial IntelligenceOct 6

Musubi Launches Open-Weight PolicyLM-1.7B for Real-Time AI Moderation