What happened
OpenAI previewed a new API service tier called Ultrafast. The tier runs the GPT-5.6 Sol model up to 14 times faster than before. It processes up to 750 output tokens per second. Cerebras hardware powers the speed increase.
The context
Token speed controls how quickly an AI model writes text or code back to a user. Faster output reduces latency for interactive applications. Using custom hardware helps OpenAI push real-time output rates higher for developers.
Sources
- Official source, https://openai.com/index/previewing-ultrafast Primary source