
OpenAI announced today (August 14) the preview release of Ultrafast mode for its most powerful AI model, GPT-5.6 Sol. The mode is 14 times faster than standard processing, with output of up to 750 tokens per second.
For service support, OpenAI said the mode is provided by Cerebras, with the goal of further improving model response speeds in users' workflows.

OpenAI said that in the past, model speed and capability could not both be maximized. Users who needed faster output generally had to sacrifice model capability and switch to smaller or more specialized models.
Ultrafast makes it possible to have both, establishing “more valuable work completed per second” as a new direction. OpenAI's early use cases include incident response and reliability analysis, financial research and security analysis, customer service and voice interaction, product Q&A and inventory inquiries in e-commerce, as well as real-time research and experimentation. Relevant screenshots are shown below:

OpenAI said the mode is intended for business scenarios with extremely high real-time requirements. It is currently in limited preview and available only to a small group of customers. The company said these early tests will help determine where speed improvements create the most value and assess whether the model can keep pace with users' actions.
