OpenAI has unveiled a new inference mode called Ultrafast, promising a dramatic leap in processing speed for its GPT-5.6 Sol model. This enhancement, powered by Cerebras hardware, delivers up to 750 output tokens per second, marking a 14x improvement over previous performance levels.
Speed as a Service
The launch introduces a three-tiered pricing model for inference: Standard, Fast, and now Ultrafast. This move positions speed not just as a performance feature, but as a distinct product offering, enabling developers and enterprises to select the processing power that best suits their needs.
The collaboration between OpenAI and Cerebras, under their $10 billion partnership, underscores a growing industry trend where specialized hardware is becoming essential for optimizing AI workloads. By leveraging Cerebras’ high-performance chips, OpenAI is able to meet the increasing demand for rapid AI responses in real-time applications.
Implications for Developers and Enterprises
This development is particularly significant for businesses that rely on large language models for customer service, content generation, and real-time analytics. The Ultrafast mode could reduce latency and improve user experience, especially in interactive applications where response time is critical.
Industry analysts suggest this move may further solidify OpenAI’s dominance in the AI inference space, especially as competitors race to offer similar performance improvements. The ability to scale inference speed could also drive innovation in AI-powered tools, enabling more dynamic and responsive applications.
Conclusion
With the introduction of Ultrafast mode, OpenAI is not only enhancing the capabilities of its models but also reshaping how inference speed is perceived and monetized. As AI systems become more embedded in everyday applications, such advancements could be pivotal in defining the next generation of AI-powered services.



