OpenAI Ultrafast Makes GPT-5.6 Sol 14x Faster With Cerebras
OpenAI’s Ultrafast preview runs GPT-5.6 Sol at up to 14 times standard speed and 750 tokens per second through its Cerebras chip partnership for enterprises.
Summary
OpenAI introduced Ultrafast in preview on August 13, 2026, a mode it says runs GPT-5.6 Sol, its latest and most powerful model, at 14 times standard processing speed and generates up to 750 output tokens per second. Output tokens are the distinct text units an LLM produces during an interaction. OpenAI says Ultrafast seeks real-time performance without requiring a smaller or specialized model, delivering more useful work per second.
OpenAI targets corporate incident response, customer service and support, financial market analysis, e-commerce and other workflows. Cerebras powers Ultrafast through its chip partnership with OpenAI. The preview is limited to a small customer group, with access due to expand as capacity grows. Anthropic also accelerates Claude through fast mode, although TechCrunch reports it does not reach OpenAI’s stated speed.
Positives
- Ultrafast delivers 14 times GPT-5.6 Sol’s standard processing speed, according to OpenAI.
- Up to 750 output tokens per second could support workflows requiring rapid model responses.
- Cerebras powers Ultrafast through its chip partnership with OpenAI.
- Incident response, customer support, financial analysis and e-commerce are among OpenAI’s proposed corporate applications.
- GPT-5.6 Sol gains real-time speed without switching to a smaller or specialized model, OpenAI says.
Risks & concerns
- Only a small group of customers can access the Ultrafast preview.
- Broader availability depends on capacity growth, and OpenAI provided no expansion timetable.
- Ultrafast remains a preview rather than a generally available GPT-5.6 Sol feature.
- The reported 14 times speed and 750-token rate are OpenAI’s claims, with no independent benchmark cited.