Thursday, August 27, 2026
Tech Beat

Meta Challenges Claude Code and Codex With Muse Code and Muse Spark 1.2

Meta launches Muse Code and Muse Spark 1.2, pairing persistent AI agents with low-cost tokens, strong benchmarks and significant code privacy tradeoffs.

A hollow circuit-shaped key unlocks a secure vault while secretly drawing glowing data fragments away.
Listen to this briefingAudio briefing

Summary

Verified facts, Meta released Muse Code in beta on August 5, 2026, alongside Muse Spark 1.2, a proprietary model optimized for software development. The terminal-based agent runs on macOS and Linux and is designed to plan, implement and validate changes across large repositories. Installation uses a single shell command, but developers must sign in with a Meta account and supply billing information. The launch puts Meta into more direct competition with Anthropic’s Claude Code, OpenAI’s Codex and specialist coding platforms such as Cursor.

Verified facts, Muse Code’s main architectural distinction is a collection of specialized background agents that remain active for an entire session instead of being recreated for every assignment. Meta argues that this persistence reduces repeated repository analysis and allows agents to continue work with less supervision. Large jobs can be divided among parallel sub-agents operating in isolated Git worktrees; Mark Zuckerberg said an internal test produced six game features simultaneously without collisions. An append-only local log records model calls, approvals, tool activity and edits before execution. Meta says this design allows an interrupted job, including one running for 20 hours, to restart from the same state. Those resilience claims have not yet been independently validated at scale.

Verified facts, Muse Spark 1.2 was co-trained with the Muse Code harness and used synthetic coding environments and grading generated by Muse Spark 1.1. In Meta’s published results, version 1.2 scored 82.9% on Terminal-Bench 2.1, ahead of GPT-5.6 Terra at 81.8% but behind Claude Opus 5 at 86.7%. It reached 59.3% on DeepSWE 1.1, trailing Opus 5 at 65.0% and GPT-5.6 Terra at 64.8%. On Meta’s internal coding test, it scored 70.6% versus 79.4% for Opus 5. Meta reported gains of 6.7 points on Terminal-Bench and 6.3 points on DeepSWE over Muse Spark 1.1, although the older model used a different harness, preventing a clean model-only comparison.

Verified facts, Meta also demonstrated the agent optimizing GPU kernels on NVIDIA Hopper hardware over as many as 24 hours and more than 1,000 tool calls. The system wrote and profiled Triton code without wrapping existing third-party kernel libraries, and Meta reported improvements to KDA and MLA kernel baselines. This is a company-produced case study, so it remains uncertain whether comparable progress will occur on external repositories. Pricing is split between a standard API tier at $1.25 per million input tokens and $4.25 per million output tokens, and a contributor tier at $0.10 and $0.20 respectively. Standard-tier data is excluded from model training; contributor customers explicitly permit Meta to train on their prompts and completions.

Verified context, The contributor plan is limited to 60 requests per minute, compared with 3,000 on the standard tier, and still requires a payment method. VentureBeat installed a 97 MB package on a Mac mini but could not run the agent until payment setup was completed. The proprietary release also marks a break from Meta’s former Llama strategy. Llama models had accumulated about 1.2 billion downloads by early 2026, but Muse Spark and Muse Code provide neither downloadable weights nor self-hosting. Zuckerberg said he would share more about possible open sourcing soon, without making a specific commitment.

Interpretation and outlook, Muse Code gives Meta a credible entry into an important enterprise AI market through persistent agents, recoverable execution and unusually aggressive pricing. Developers, independent teams and engineering leaders could benefit from lower costs and more auditable long-running jobs. The central risks are whether real-world performance matches Meta’s demonstrations and whether organizations will accept sending proprietary source code into Meta’s training pipeline for a discount. Claude Opus 5 still led all three cited benchmark charts, so Meta has not established overall technical leadership. Adoption, enterprise trust, independent evaluations and any future open-source announcement will determine whether the beta can materially disrupt Anthropic and OpenAI.

Positives

  • Muse Spark 1.2 scored 82.9% on Terminal-Bench 2.1, exceeding GPT-5.6 Terra’s 81.8% result in the published comparison.
  • Meta’s append-only local event log is designed to preserve every model call, approval, tool action and edit, enabling interrupted long-running tasks to resume without starting over.
  • Muse Code can divide large assignments among isolated Git worktrees, and Zuckerberg reported that an internal test built six game features in parallel without collisions.
  • The contributor tier costs $0.10 per million input tokens and $0.20 per million output tokens, substantially reducing the price of experimentation for users who accept training-data use.
  • Muse Spark 1.2 improved by 6.7 points on Terminal-Bench and 6.3 points on DeepSWE compared with Meta’s reported results for version 1.1.

Risks & concerns

  • Claude Opus 5 beat Muse Spark 1.2 on all three cited evaluations, including an 86.7% versus 82.9% lead on Terminal-Bench 2.1 and 79.4% versus 70.6% on Meta’s internal benchmark.
  • Customers using the discounted contributor tier give Meta permission to train future models on their prompts and completions, creating a serious concern for proprietary or regulated code.
  • Muse Code requires a Meta account, billing details and a payment method even on the contributor tier, as VentureBeat discovered after installing the 97 MB client.
  • The claimed generational improvement is not a model-only comparison because Muse Spark 1.1 ran in the generic mini-swe-agent harness while version 1.2 ran in Muse Code.
  • Muse Code and Muse Spark 1.2 are proprietary, with no downloadable weights, self-hosting option or confirmed open-source release despite Meta’s earlier Llama strategy.
  • Meta’s 24-hour GPU optimization demonstration has not yet been independently reproduced, leaving its applicability to outside repositories uncertain.
Primary sourceVentureBeathttps://venturebeat.com/orchestration/meta-enters-the-ai-coding-wars-with-muse-spark-1-2-and-muse-code-with-persistent-async-background-agents
Read full article
Editorial note: Tech Beat summarizes and analyzes third-party reporting. The source link is the authoritative article. This page does not reproduce the full source text.

More From The Wire

CybersecurityAug 27

Visa VVAH AI Patches Code Before Human Review

Artificial IntelligenceAug 27

OpenAI Brings ChatGPT Ads to India With 50 Brands, ₹725 Daily Floor

Artificial IntelligenceAug 27

Nvidia Nears $12.9 Billion Hugging Face Acquisition Amid Conflicting Reports