SpaceXAI Grok 4.6 Ties GPT-5.6 Sol Max With Lower API Pricing
SpaceXAI's Grok 4.6 ties GPT-5.6 Sol Max at 61, boosts agent and coding scores, and starts at $2 input and $6 output per million tokens for enterprises.
Summary
On August 12, 2026, Elon Musk's SpaceXAI, renamed after SpaceX acquired xAI in February, released Grok 4.6 for long-running coding and knowledge agents. It follows July's Grok 4.5 by weeks and Grok Bot by one day. The model is available through the API, Grok Build from the $30 monthly SuperGrok plan, acquired startup Cursor, OpenRouter, Vercel and Cloudflare, with double included usage in Cursor and Grok Build for one week. Its 500,000-token context supports text and image inputs, reasoning, function calling and structured text output.
Artificial Analysis scores Grok 4.6 at 61, five points above Grok 4.5 High, ahead of Moonshot AI's open-weights Kimi K3 and tied with OpenAI's GPT-5.6 Sol Max, behind Anthropic's Claude Opus 5 and Fable 5. GDPVal-AA v2 rose from 1,526 to 1,753 Elo, topping GPT's 1,728 and Fable's 1,741. Scores reached 69.9% on CursorBench, 65.9% DeepSWE, 61.3% FrontierCode, 57.5% APEX-Agents, 56.4% APEX-SWE and 26% Terminal-Bench, versus frontier leaders of 70.5%, 73%, 63.6%, 59.2%, 58.8% and 34.6%, respectively. Grok led AA-Briefcase at 1,577 Elo and Harvey LAB at 15.8%, but mixed self-reported and public results prevent controlled comparison.
API pricing is $2 per million input and $6 output tokens, $8 combined versus $35 for GPT-5.6 Sol standard, while a faster version doubles rates. At 200,000 prompt tokens, all rates rise to $4 input, $1 cached and $12 output. Artificial Analysis measured $0.84 per task, 53 turns and 0.5 billion input tokens, versus Claude Opus 5 Max's 103 turns and 2 billion. Adoption faces Grok's 2025 antisemitic, extremist, political and Musk-flattery incidents and 2026 sexual-imagery investigations by Ofcom, Britain's Information Commissioner's Office and the European Commission, although none establishes Grok 4.6 violated law.
Positives
- Grok 4.6 scored 61 on Artificial Analysis, beating Kimi K3, gaining five points over Grok 4.5 High and tying GPT-5.6 Sol Max.
- GDPVal-AA v2 climbed from 1,526 to 1,753 Elo, surpassing GPT-5.6 Sol Max at 1,728 and Fable 5 Max at 1,741.
- APEX-Agents improved 10.4 points to 57.5%, narrowly beating GPT-5.6 Sol Max's 56.7%.
- AA-Briefcase averaged 53 turns and 0.5 billion input tokens, compared with Claude Opus 5 Max's 103 turns and 2 billion.
- $2 input and $6 output pricing per million tokens makes standard Grok 4.6 less than half GPT-5.6 Sol standard's $35 combined rate.
- Cursor, Grok Build, OpenRouter, Vercel and Cloudflare provide immediate distribution, with double included Cursor and Grok Build usage during the first week.
Risks & concerns
- Terminal-Bench reached only 26%, trailing GPT-5.6 Sol Max at 34.6% and Fable 5 Max at 34.1%.
- $0.84 per task made Grok 4.6 less economical than Grok 4.5, GPT-5.6 Luna, GLM-5.2 and Muse Spark 1.2.
- Prompts reaching 200,000 tokens double input, cached-input and output rates across every token in the request.
- Mixed self-reported and publicly available benchmark results prevent a fully controlled comparison with GPT-5.6 Sol Max and Fable 5 Max.
- Ofcom, Britain's Information Commissioner's Office and the European Commission continue investigations connected to earlier Grok sexualized-image generation and X's safeguards.
- Grok's 2025 antisemitic, extremist, political and Musk-flattery outputs create governance, compliance and brand-safety risks for regulated enterprise buyers.