By Daily Touch Insights Editorial Team
Editorial Team
View Journalist Profile
AI & TECHNOLOGY — OpenAI is giving its most capable GPT-5.6 Sol model a major speed boost with a new Ultrafast mode that can generate up to 750 output tokens per second, potentially making advanced AI practical for applications where even small delays matter.
The new service tier is being introduced as a limited preview through the OpenAI API. OpenAI says GPT-5.6 Sol on Ultrafast can run up to 14 times faster than its Standard processing mode. 0
What Is GPT-5.6 Sol Ultrafast?
Ultrafast is not a smaller or less capable AI model.
It is a faster processing tier for GPT-5.6 Sol, OpenAI's flagship model in the GPT-5.6 family.
The idea is straightforward: keep the intelligence of the full model while dramatically reducing the time required to generate an answer.
OpenAI says the system can reach speeds of up to 750 output tokens per second, compared with Standard processing, which makes the maximum claimed improvement as high as 14×. 1
Cerebras Is Powering the Speed Boost
The acceleration comes from Cerebras, a company that develops specialized AI computing hardware.
Cerebras' architecture is designed around wafer-scale processors that can provide extremely high-speed access to the data required during AI inference.
That architecture is particularly useful for reducing latency when a model is generating responses token by token.
OpenAI says Ultrafast is powered by Cerebras as part of the companies' collaboration on low-latency AI inference. 2
Why Speed Matters for Advanced AI
AI capability is only useful if people can interact with it efficiently.
A model may be extremely intelligent, but if every response takes too long, it becomes difficult to use in situations where humans expect immediate feedback.
That limitation becomes even more important when AI agents need to perform multiple steps.
An agent might need to read information, make a decision, use a tool, inspect the result and then make another decision.
If every step introduces significant waiting time, the entire workflow becomes slow.
Ultrafast is designed to attack that problem.
OpenAI Is Targeting Real-Time Work
OpenAI says early Ultrafast customers are testing the technology across areas including coding, commerce, financial research, customer support and security response.
These are environments where response time can directly affect productivity.
A developer waiting for an AI coding agent to complete a task, for example, could potentially receive results much faster and immediately continue to the next stage.
Voice AI Could Benefit Significantly
Real-time voice interaction is another obvious application.
People naturally expect conversations to move quickly.
Long pauses between a question and an AI response can make a system feel artificial and frustrating.
Higher inference speeds could allow more sophisticated models to participate in conversations without requiring companies to sacrifice intelligence simply to reduce latency.
OpenAI specifically identifies real-time voice and other live applications as potential uses for Ultrafast. 3
AI Agents Could Become Much More Practical
The biggest impact may not be ordinary chat.
It could be autonomous AI agents.
Agents often need to perform long sequences of actions. Faster model responses mean those sequences can potentially be completed in much less time.
For businesses, that could make AI systems more useful for research, software development, financial analysis and customer operations.
The advantage compounds when a workflow requires dozens or hundreds of model interactions.
OpenAI Says Intelligence Does Not Have to Be Sacrificed
Historically, companies have often faced a trade-off between intelligence and speed.
Smaller models can respond quickly, while larger frontier models generally require more computing resources.
Ultrafast is an attempt to change that equation by running the full GPT-5.6 Sol model on specialized hardware.
Cerebras says the Ultrafast version maintains the model's capabilities while delivering substantially lower latency. 4
It Is Still Only a Limited Preview
This is where the headline needs some caution.
Ultrafast is not currently a universally available feature.
OpenAI is initially making it available to a select group of customers through the API while it evaluates the technology and expands capacity.
Broader access is expected to depend on capacity and workload requirements. 5
Therefore, saying GPT-5.6 Sol is now 14 times faster for everyone would be misleading.
750 Tokens Per Second Is a Maximum
The phrase “up to 750 tokens per second” is also important.
Maximum performance is not necessarily the speed every customer will experience on every request.
Actual latency can depend on factors including workload complexity, request size, system conditions and infrastructure.
The 14× figure should therefore be understood as OpenAI's maximum stated improvement rather than a guarantee that every response will be exactly 14 times faster.
GPT-5.6 Is Already Designed for Complex Work
OpenAI's GPT-5.6 family was introduced with a focus on coding, knowledge work, scientific reasoning, cybersecurity and agentic workflows.
GPT-5.6 Sol is positioned as the flagship model, while the family also includes GPT-5.6 Terra and GPT-5.6 Luna for different performance and cost requirements. 6
Ultrafast therefore adds another dimension to the model family: not simply intelligence or cost, but response speed.
The Combination Could Change AI Workflows
Consider a financial research agent that needs to analyse hundreds of documents.
The model might repeatedly search, read information, compare evidence and produce conclusions.
If every interaction becomes faster, the agent can potentially complete the overall workflow much sooner.
The same principle applies to software engineering, customer support and complex research.
Speed becomes especially valuable when AI is doing many things rather than answering one question.
Companies May Build Different Products Around It
Ultrafast could also change how developers design AI products.
Instead of building applications around slow, asynchronous AI interactions, developers could create systems that respond almost continuously as users interact with them.
This could enable more responsive AI tutors, coding assistants, research systems, shopping agents and enterprise support tools.
The product experience could begin to resemble traditional software rather than a chatbot that users periodically wait for.
There Is Still a Cost Question
Extreme inference speed is unlikely to be free.
Running advanced models at very high throughput requires substantial computing infrastructure.
For businesses, the important question will not simply be how fast the model is.
It will be whether the additional speed produces enough economic value to justify the cost.
A company processing millions of customer interactions may find the economics attractive, while a casual user may not need that level of performance.
Cerebras Gains a Major Opportunity
The partnership also puts Cerebras in an important position within the AI infrastructure market.
The company has long argued that its wafer-scale approach can deliver extremely fast AI inference.
Running a leading frontier model such as GPT-5.6 Sol on its hardware gives that technology a high-profile real-world test.
If customers find the performance compelling, the partnership could strengthen Cerebras' position as AI companies increasingly look beyond traditional GPU infrastructure.
The AI Infrastructure Race Is Changing
The competition in artificial intelligence is no longer only about who builds the smartest model.
It is increasingly about who can deliver intelligence at the lowest cost and lowest latency.
Companies are competing across several layers: models, chips, data centres, networking, software and inference systems.
Ultrafast illustrates how important that infrastructure competition has become.
Faster AI Could Create New Problems Too
Greater speed is not automatically better in every situation.
If an AI system can generate large amounts of information almost instantly, users may receive incorrect or poorly considered outputs more quickly as well.
For high-stakes applications, accuracy, verification and safety remain important even when latency falls dramatically.
Speed should therefore complement intelligence and reliability, not replace them.
What Happens Next?
OpenAI plans to expand access to Ultrafast as capacity grows.
The company is also using the preview period to learn where a major increase in inference speed creates the greatest value.
That feedback could influence how future AI products are designed.
If customers consistently prefer highly responsive frontier models, low-latency inference could become a standard expectation rather than a premium feature.
Our Perspective
The important part of Ultrafast is not simply the impressive “14×” headline.
The bigger development is the attempt to remove one of the fundamental limitations of frontier AI: waiting.
When powerful models become fast enough to respond almost as quickly as people interact with software, entirely different product experiences become possible.
But the real test will come from actual customer workloads, not laboratory maximums.
If OpenAI can make frontier-level intelligence consistently fast, economically viable and widely available, AI could shift from something people wait for into something that works alongside them continuously.
Conclusion
OpenAI is previewing Ultrafast, a new service tier that runs GPT-5.6 Sol at speeds of up to 750 output tokens per second—up to 14 times faster than Standard processing. The technology is powered by Cerebras and is initially available to a limited group of API customers. 7
The technology is aimed at workloads where response time matters, including coding, customer support, commerce, financial research, security and real-time voice applications.
However, the 14× figure represents maximum claimed performance rather than a universal guarantee, and access remains limited during the preview.
The real significance of Ultrafast is simple: if frontier AI becomes fast enough to operate at the speed of human interaction, the way people use intelligent software could change dramatically.
