OpenAI and Cerebras Introduce Ultrafast

On August 13, 2026, OpenAI and Cerebras Systems announced Ultrafast, a new inference mode built for GPT-5.6 Sol. Per OpenAI's official announcement, the mode lets the model generate responses at up to 750 output tokens per second, a jump the company describes as up to 14 times faster than the same model's standard processing. The partnership with Cerebras, a company known for hardware built for accelerated AI processing, underpins this extra layer of performance inside the API.

What Actually Changes Under the Hood

Ultrafast is not a new model. It is an additional mode, or tier, of access to the existing GPT-5.6 Sol, offered through the API. In practice, developers keep the same model and the same response quality, but switch the processing path to prioritize token generation speed. As TechCrunch reported, the performance gain is credited directly to Cerebras's infrastructure, which confirms that the difference sits in the inference layer, not in the model itself.

Why 750 Tokens Per Second Matters

In asynchronous text applications, generation speed often goes unnoticed by end users. In real time interactions, such as support chatbots or voice assistants, every hundredth of a second of delay is noticeable and directly shapes the experience. Running at 750 tokens per second sharply cuts the gap between a user's question and the start of the answer, which matters most in spoken conversations, where long pauses sound unnatural and break the flow of dialogue.

Use Cases Tested in the Preview

According to Help Net Security and OpenAI's own announcement, early Ultrafast testing focused on five areas: real time customer service, conversational commerce, incident response, real time financial analysis and real time voice interactions. These are precisely the scenarios where latency has traditionally limited production adoption of AI agents, since any noticeable delay chips away at user trust in the tool.

Limited Access: Not Yet Generally Available

This point deserves emphasis: Ultrafast is in limited preview, available only to select customers through the API. OpenAI has not published a confirmed date for general availability, meaning most developers and businesses do not yet have direct access to the mode. Both TechCrunch and Help Net Security frame the launch as an early technology preview rather than a fully available product.

No Public Pricing Yet: A Reason for Caution

As of this announcement, OpenAI has not published pricing for using Ultrafast through the API. High speed inference tiers have historically carried a higher per token cost than standard processing, but any specific figure at this stage would be pure speculation. For agencies and businesses considering offering this capability to clients, the absence of a public price sheet is a clear signal that now is not the time to make commercial promises around it.

Coverage Across the Tech Press

The announcement drew quick coverage from outlets including TechCrunch, Help Net Security, 9to5Mac and Tech Times, alongside Cerebras's own statement to investors confirming its partnership with OpenAI to power the new mode's infrastructure. The overlap between independent tech reporting and a corporate investor release confirms this is a real, coordinated announcement between both companies, even as commercial and availability details remain open.

Why It Matters for Brazilian Agencies and SMBs

For marketing agencies and support teams already using or evaluating AI agents, Ultrafast signals where the industry is heading: less emphasis on raw model intelligence gains, more emphasis on response speed, especially for real time voice and chat support. Still, because the feature sits in a closed preview with no public pricing, the sensible move right now is to track the rollout without folding Ultrafast into client proposals or deadlines. Once general availability and pricing are confirmed, cutting latency could become a real differentiator for SMBs that depend on fast automated support, but promising it too early risks setting expectations OpenAI cannot yet back up.