Microsoft ThinkingBox Exposes the Reliability Gap in Enterprise AI Agents
Microsoft ThinkingBox tests 507 enterprise AI workflows 20 times, exposing clean-looking runs that leave wrong database states and costly reliability gaps.
Summary
Microsoft released ThinkingBox on Hugging Face through OpenEnv on October 3, 2026. ThinkingBox-Bench runs 507 synthetic retail, auto insurance, travel, neobank and consulting workflows 20 times from clean MCP sessions, grading terminal database state and side effects. In one $745 appliance case, nine plausible tool calls closed a ticket as solved although a Nashville courier exception remained open 15 days late, requiring hold status.
A 12-model ablation covered 121,680 valid trials: 79,853 failed executable checks, while 67.24% of failures ended cleanly after a state-changing call with no final tool error. Checks found wrong fields in 77.61%, extra effects in 43.30% and missing effects in 25.36%, with overlap. Among 18 leaderboard models, Claude Opus 5.5 led pass@1 at 67.16%. Kimi-K3 solved 93.89% of tasks at least once, but only 68, or 13.41%, passed 20/20. Claude Opus 5 and 5.5 each achieved 241 dependable tasks, or 47.53%, despite pass@1 scores of 66.50% and 67.16%.
Tool usage represented 79.9% of failures, followed by wrong state updates at 10.3%, incomplete resolutions at 7.0% and no state-changing action at 2.9%. GPT-5.4 was cheapest per dependable task at $6.80 for 128 tasks, versus GPT-6 Astra at $7.45 for 231 and Claude Opus 5.5 at $7.80 for 241. Of 507 tasks, 477 use state-only grading and 30 add response rubrics. The MIT-licensed harness and CDLA-Permissive-2.0 dataset are available now, with reinforcement learning training planned.
Positives
- Claude Opus 5.5 led the 18-model leaderboard with 67.16% pass@1 and completed 241 tasks correctly in all 20 trials.
- Kimi-K3 solved 476 of 507 tasks at least once and led retail workflows with 82.24% pass@1.
- GPT-5.4 delivered the lowest estimated cost per dependable task, $6.80 across 128 consistently completed workflows.
- 477 of 507 tasks use deterministic state grading, reducing dependence on subjective evaluations of agent responses.
- ThinkingBox and ThinkingBox-Bench are publicly available through Hugging Face and OpenEnv for reproducible evaluation.
Risks & concerns
- 79,853 of 121,680 valid trials failed executable checks, despite 67.24% of failures ending cleanly after changing state without a final tool error.
- 77.61% of failed trials contained wrong field values, while 43.30% created extra effects and 25.36% omitted required effects.
- Kimi-K3 reached 93.89% of tasks at least once, yet only 13.41% succeeded in all 20 attempts.
- Claude Opus 5.5 improved pass@1 over Opus 5 but added no dependable tasks, with both completing 241 workflows 20 out of 20 times.
- 79.9% of failures involved tool handling, highlighting weak recovery from tool errors, failed preconditions and empty lookups.
- Auto insurance averaged 33.83% pass@1 across listed models, far below retail's 59.52%, showing sharp domain sensitivity.