Skip to main content
Product Updates

Up to 750 Tokens per Second: OpenAI Introduces Ultrafast Mode, Boosting GPT-5.6 Sol AI Speed 14x

OpenAI announced today (August 14) the preview release of Ultrafast mode for its most powerful AI model, GPT-5.6 Sol. It delivers 14 times the standard processing speed, with output of up to 750 tokens per second.

Up to 750 Tokens per Second: OpenAI Introduces Ultrafast Mode, Boosting GPT-5.6 Sol AI Speed 14x

OpenAI announced today (August 14) the preview release of Ultrafast mode for its most powerful AI model, GPT-5.6 Sol. The mode is 14 times faster than standard processing, with output of up to 750 tokens per second.

For service support, OpenAI said the mode is provided by Cerebras, with the goal of further improving model response speeds in users' workflows.

Up to 750 Tokens per Second: OpenAI Introduces Ultrafast Mode, Boosting GPT-5.6 Sol AI Speed 14x

OpenAI said that in the past, model speed and capability could not both be maximized. Users who needed faster output generally had to sacrifice model capability and switch to smaller or more specialized models.

Ultrafast makes it possible to have both, establishing “more valuable work completed per second” as a new direction. OpenAI's early use cases include incident response and reliability analysis, financial research and security analysis, customer service and voice interaction, product Q&A and inventory inquiries in e-commerce, as well as real-time research and experimentation. Relevant screenshots are shown below:

Up to 750 Tokens per Second: OpenAI Introduces Ultrafast Mode, Boosting GPT-5.6 Sol AI Speed 14x

OpenAI said the mode is intended for business scenarios with extremely high real-time requirements. It is currently in limited preview and available only to a small group of customers. The company said these early tests will help determine where speed improvements create the most value and assess whether the model can keep pace with users' actions.