OpenAI Launches Ultrafast Mode for GPT-5.6 Sol
OpenAI is pairing its frontier model with Cerebras hardware to push GPT-5.6 Sol to 750 tokens per second, a speed tier that changes what real-time use of the model can look like.
Reporting from 1 source: GIGAZINE.
OpenAI announced Ultrafast, a service that runs GPT-5.6 Sol up to 14 times faster, using Cerebras inference technology. The preview is open to some users, with a waitlist for broader access. OpenAI says Ultrafast reaches 750 tokens per second and completes Humanity's Last Exam in 11 hours 11 minutes.
OpenAI already offered Fast mode, which runs GPT-5.6 Sol at up to 2.5 times its standard speed. Ultrafast goes further, claiming up to 14 times the speed. The company says the mode sustains state-of-the-art performance while producing 750 tokens per second.
On the 2,500-question benchmark Humanity's Last Exam, Ultrafast finished in 11 hours 11 minutes. OpenAI also cites its GDP-Val test, where Ultrafast handled practical work at 5.6 times the speed of standard mode.
The preview is available to a limited set of users. OpenAI plans to expand access as compute becomes available, and interested users can join a waitlist.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.