Holo4 Open-Weight AI Agents Control GUIs, Code, MCP and APIs
Holo4 brings GUI, code, MCP and API control to two open-weight agents, challenging frontier models on cost while trailing Opus 5.5 in OSWorld 2.0 testing.
Summary
H launched Holo4 on September 28, 2026, as 27B dense and 35B-A3B Mixture of Experts models, alongside Holotron4 Nano. One model operates across desktop, web, Android, code sandboxes and business systems by clicking GUIs, writing and running code, or calling MCP and API tools. On OSWorld 2.0, Holo4 27B scored 61.7%, behind Opus 5.5 at 81.8%, while Holo4 35B-A3B reached 30.9%. H claims substantially lower task costs and orders of magnitude fewer parameters than leading closed models.
Supervised and reinforcement learning used about 10,000 tasks generated by H’s Agentic Task Factory across web apps, MCP servers, desktops and hybrid environments. A rebuilt execution harness added memory spanning hundreds of steps and a desktop shell. In an autonomous Godot game task, Holo4 27B used 68 calls and 2.4 million tokens, versus Qwen3.8 27B’s 197 calls and 11.4 million. Benchmark comparisons use differing releases, subsets, harnesses and pricing assumptions, and Holo4 still awaits AutomationBench private-set evaluation. Both Holo4 sizes are available through the H Models API, with BF16, FP8, NVFP4 and 4-bit GGUF weights, plus reproducible trajectories, on Hugging Face. Optimized DSpark drafter checkpoints are due in the coming days.
Positives
- Holo4 27B scored 61.7% on OSWorld 2.0 while using far fewer parameters and lower claimed task costs than leading closed models.
- One Holo4 model works across desktop, web, Android, code sandboxes, GUIs, MCP tools and business APIs.
- About 10,000 generated tasks trained Holo4 across web apps, desktop software, MCP servers and hybrid environments.
- Holo4 completed the Godot task in 68 calls and 2.4 million tokens, versus Qwen3.8 27B’s 197 calls and 11.4 million.
- BF16, FP8, NVFP4 and 4-bit GGUF weights and public benchmark trajectories support deployment and reproducibility.
Risks & concerns
- Holo4 27B’s 61.7% OSWorld 2.0 score trails Opus 5.5’s 81.8% by 20.1 percentage points.
- Holo4 35B-A3B reached only 30.9% on OSWorld 2.0, substantially below the stronger 27B model.
- Benchmark comparisons involve different releases, subsets, harnesses and pricing assumptions, limiting direct cost and performance conclusions.
- AutomationBench results use H’s internal harness and public-set testing, with private-set evaluation still pending.
- Professional workflow examples consumed up to 2.4 million tokens, showing that complex autonomous tasks can remain computationally intensive.