OpenAI has announced a preview of Ultrafast, a new service tier for its API that dramatically accelerates inference for its latest model, GPT-5.6 Sol. The company claims the tier delivers up to 14× the speed of standard inference, with throughput reaching as high as 750 output tokens per second.
The speed boost is powered by Cerebras, a company known for its wafer-scale AI chips. This partnership marks a significant step in making large language models more responsive for real-time applications, such as interactive chatbots, coding assistants, and high-frequency content generation.
Ultrafast is currently in preview, and OpenAI has not yet disclosed pricing or general availability details. Developers interested in testing the new tier can likely expect integration through the existing OpenAI API infrastructure.
This move aligns with the industry trend toward reducing latency and increasing token throughput, which is critical for scaling AI applications that require near-instantaneous responses.