The short version
OpenAI launches Ultrafast mode for GPT-5.6 Sol, delivering a 14x speed boost to 750 tokens/sec via a partnership with Cerebras.
OpenAI has previewed a new ‘Ultrafast’ mode for its flagship GPT-5.6 Sol model, powered by a partnership with Cerebras. This mode reaches up to 750 output tokens per second, a 14x speed gain. The service is first available through the OpenAI API to a chosen set of customers.
Key takeaways
- OpenAI’s new Ultrafast mode for GPT-5.6 Sol delivers up to 750 tokens per second, a 14x speed increase.
- The acceleration is powered by a partnership with Cerebras, stemming from a multi-billion dollar deal.
- The mode enables real-time applications in areas like incident response, finance, and customer support.
- Access is initially limited to a select group via the OpenAI API, with plans for gradual expansion.
- The launch introduces a new, faster tier to OpenAI’s existing tiered pricing model for inference speed.
The Ultrafast Launch: Partnership, Performance, and Availability
OpenAI is launching a preview of its new “Ultrafast” mode for its flagship model, GPT-5.6 Sol. This mode delivers up to 750 output tokens per second, which represents a 14x speed increase.
The inference acceleration for this new mode comes from a partnership with Cerebras. This collaboration stems from a ten-billion-dollar deal signed between the two companies earlier this year.
Limited Initial Rollout
The Ultrafast service will first be available only through the OpenAI API for the GPT-5.6 Sol model. Access is limited to a chosen set of customers at launch. OpenAI states it plans to expand availability as its capacity to support the mode increases.
Enabling New Real-Time Use Cases Across Industries
OpenAI states that its new Ultrafast mode merges the speed of smaller models with the full abilities of a large reasoning model. This allows what the company calls “more useful work per second.” The speed gain creates several real-time applications across different sectors.
Incident Response and Finance
In incident response, engineers could have logs, code changes, and reports analyzed while an outage is still happening. This helps them find the cause and prepare a fix immediately. OpenAI says it is already using the model internally for this purpose. In finance, the model could evaluate shifting market signals and flag suspicious transactions as conditions change.
Customer Support and E-commerce
For customer support, complex, multi-step inquiries could be resolved instantly. In e-commerce, the model could answer product questions, check inventory levels, and personalize recommendations before a hesitant buyer leaves their cart.
Research
Research is another key area. Experiments that previously ran overnight as batch jobs could become interactive work sessions. Teams could test an idea, review results, adjust their approach, and start another run without breaking their workflow.
Business Model: Monetizing Speed with Tiered Pricing
OpenAI already monetizes inference speed through tiers. Its API offers a “Fast Mode” for GPT-5.6 Sol that promises up to 2.5x speed at roughly double the price. The new Ultrafast mode introduces a third, faster, and likely pricier tier to this existing structure.
This pricing logic directly mirrors the strategy of cloud providers like AWS, which have long charged more for higher performance levels of the same core service. By applying this model to AI inference, OpenAI positions itself to capture a share of the revenue gains that faster inference creates for its customers. The company argues that as speed becomes a critical bottleneck across industries—from real-time finance analysis to interactive customer support—its tiered pricing allows it to benefit directly from the increased utility and efficiency it provides.
📡 Original reporting: The Decoder. AI Craft Technologies’ news engine summarised and rewrote this story in our own words; facts are drawn from the linked source.
⚙️ How this article was made — fully automated
This is a live demo of the ACT News Factory engine. Want one running on your own site? See our services →



