By Daily Touch Insights Editorial Team
Editorial Team
View Journalist Profile
ARTIFICIAL INTELLIGENCE & TECHNOLOGY — OpenAI is rolling out a major set of performance improvements across ChatGPT and Codex, with reported gains aimed at making long conversations faster, reducing memory consumption and allowing AI agents to distribute complex tasks across multiple models.
The reported upgrades include a major improvement in the time required to open extremely large conversations, while OpenAI's Multi-Agent V2 system is designed to let a primary AI agent automatically delegate different parts of a task to specialised models.
According to performance figures reported from internal testing, an extremely large conversation containing 741 rounds of interaction saw its average opening time fall from 27.62 seconds to 1.66 seconds.
That represents a dramatic improvement in responsiveness for users working with very long AI sessions, particularly developers who use AI agents for coding, testing and other complex workflows.
ChatGPT's Performance Is Being Overhauled
The latest changes focus heavily on the performance of the ChatGPT interface.
Long conversations have traditionally created technical challenges because applications need to process, retrieve and display large amounts of historical information.
As AI becomes more capable of handling multi-step projects, conversations can become much longer than ordinary question-and-answer sessions.
A developer working on a large software project, for example, may use hundreds of interactions while asking an AI system to inspect code, make changes, run tests and analyse the results.
Making those conversations easier to reopen could therefore have a significant impact on productivity.
741-Round Conversation Opens in About 1.66 Seconds
One of the most striking figures from the reported testing involved a conversation containing 741 rounds of interaction.
The test conversation was reportedly around 231MB in size.
Before the optimisation, opening the conversation took an average of approximately 27.62 seconds.
After the changes, the reported average fell to approximately 1.66 seconds.
The improvement is particularly important because extremely long conversations are becoming more common as AI systems evolve from simple chatbots into tools capable of managing extended workflows.
Memory Usage Has Also Been Reduced
The performance improvements are not limited to loading speed.
Reported internal testing showed the application's memory growth during the extreme conversation test falling from approximately 1,030.7 MiB to 606 MiB.
That represents a substantial reduction in the amount of memory required to handle the session.
Lower memory consumption can make applications more responsive and reduce the risk of performance problems when users keep multiple demanding conversations open.
It could also be particularly useful for developers working on devices with limited available memory.
Network Requests Were Dramatically Reduced
Another major change concerns the number of network requests required when opening a large conversation.
According to the reported test figures, the number of requests fell from 894 to 16.
Reducing unnecessary network activity can help applications load information more efficiently.
It can also reduce the amount of work required by both the user's device and the servers supporting the application.
For users, the most visible result should simply be a faster and smoother experience.
Conversation History No Longer Needs to Load Everything Immediately
A key part of the optimisation is the way historical conversation data is handled.
Instead of attempting to load and render an entire massive conversation at once, the system can avoid processing information that is not immediately required by the user.
This approach reduces unnecessary work when a person opens an old conversation.
The strategy becomes increasingly important as AI conversations grow from a few dozen messages into hundreds or even thousands of interactions.
Multi-Agent V2 Changes How AI Models Work Together
At the same time, OpenAI is reportedly expanding its Multi-Agent V2 system.
The central idea is simple: instead of requiring a user to manually choose the right model for every part of a complex task, a primary agent can determine which models should handle different subtasks.
This creates a system in which multiple AI models can work as a coordinated team.
One model may handle a difficult reasoning problem while another works on a simpler task, allowing computational resources to be allocated according to the requirements of each part of the job.
The Main Agent Can Delegate Subtasks
Under the reported Multi-Agent V2 approach, the main agent can break a complex assignment into smaller pieces.
Those subtasks can then be delegated to other supported models.
The results can subsequently be brought back together so the main agent can produce a final response.
This is fundamentally different from the traditional chatbot model, where one model attempts to perform every part of a task sequentially.
Multi-agent systems are instead designed around cooperation between specialised AI processes.
Different Models Can Have Different Reasoning Intensities
Another reported feature is the ability to configure the reasoning intensity of individual sub-agents.
This could allow the system to use greater computational effort for difficult problems while assigning less expensive processing to routine tasks.
That approach has an important economic advantage.
If only a small portion of a complicated workflow requires the most powerful model, there may be little reason to use maximum reasoning resources for every step.
AI Is Moving Toward Automatic Model Selection
For years, users of advanced AI systems have had to decide which model is best for a particular task.
A user might select one model for coding, another for reasoning and another for faster everyday questions.
Multi-agent technology could gradually reduce the importance of that decision.
Instead of asking users to understand the strengths and weaknesses of multiple models, the AI system itself can determine which model should handle each part of a workflow.
This could make increasingly sophisticated AI systems easier to use.
Why This Matters for Developers
The changes could be especially significant for software developers.
Coding agents often need to maintain context over long periods while reading files, editing code, running tests and responding to errors.
A conversation can therefore become extremely large.
Faster access to those conversations means developers may spend less time waiting for sessions to load and more time working on their projects.
Lower memory usage could also make long-running development sessions more practical.
Long Conversations Are Becoming More Important
The growth of AI agents is changing what users expect from conversations.
Traditional chatbot interactions might involve a question followed by a short answer.
Agentic systems are different.
They can remain involved in a project for much longer, performing multiple actions and revisiting previous decisions.
That makes conversation performance an important part of the underlying AI experience rather than simply a cosmetic feature.
Speed Is Becoming a Competitive Advantage
As AI companies compete on intelligence and capability, speed is becoming equally important.
A powerful model that takes too long to respond can be frustrating when users are working through large workflows.
Faster interfaces can make advanced AI systems feel considerably more useful, even when the underlying model has not changed dramatically.
For businesses using AI at scale, faster processing can also potentially reduce wasted time and improve workflow efficiency.
OpenAI Is Targeting the AI Agent Era
The combination of faster conversation loading and multi-agent task delegation points toward a broader change in how AI systems are being designed.
AI is increasingly moving away from being treated as a simple question-and-answer service.
Instead, companies are building systems capable of planning tasks, delegating work, using tools and coordinating multiple processes.
That transition requires infrastructure capable of handling much longer and more complicated interactions.
Efficiency Could Become as Important as Intelligence
AI development has traditionally focused heavily on making models more capable.
But capability alone does not determine how useful an AI system is.
Cost, speed, memory consumption and reliability also matter.
A system that can achieve similar results using a mixture of powerful and lightweight models may be more practical than one that relies on its most expensive model for every task.
Multi-agent architecture could therefore become an important strategy for controlling the cost of increasingly sophisticated AI systems.
The Future Could Be an AI Team Rather Than One AI
The most important conceptual change may be the shift from thinking about AI as a single model to thinking about AI as a coordinated team.
In that model, the user provides the objective while the system decides how to divide the work.
Different agents can perform different functions before the results are combined.
This resembles the structure of a human organisation, where complex projects are divided among specialists rather than assigned entirely to one person.
There Are Still Important Questions
Despite the impressive performance figures, the reported results should not automatically be interpreted as meaning that every ChatGPT conversation will become 16 times faster.
Performance improvements depend on the specific workload, device, network conditions and type of conversation being processed.
The 741-round test represents an extreme workload designed to demonstrate the benefits of the optimisation.
For ordinary users, the improvement may be less dramatic, although long conversations should still benefit from the underlying changes.
Multi-Agent Systems Also Create New Challenges
Delegating tasks among multiple AI agents introduces its own technical challenges.
The system must determine which model is appropriate for each task and ensure that the different agents produce compatible results.
If an agent makes an error, that mistake could potentially affect subsequent parts of the workflow.
Reliable coordination, monitoring and evaluation will therefore become increasingly important as AI systems become more autonomous.
The User Experience Could Become Much Simpler
If automatic model selection works reliably, users may no longer need to understand the technical differences between AI models.
They could simply describe the objective they want to achieve.
The system would then determine how much reasoning is necessary, which models should be involved and how the individual tasks should be coordinated.
That could make advanced AI much more accessible to people who have no interest in learning the technical details behind the models.
What This Means for the AI Industry
The reported OpenAI changes reflect a wider industry movement toward agentic AI.
Companies are increasingly competing not only to build smarter models but also to create systems capable of completing entire workflows.
The next stage of competition could therefore focus heavily on orchestration, speed, cost efficiency and reliability.
The companies that solve those problems effectively could have a major advantage as AI becomes embedded into everyday professional work.
What Comes Next
The development of faster interfaces and multi-agent systems suggests that AI products will continue becoming more automated.
Future systems may increasingly decide which models to use, how much computing power to allocate and which tasks should be performed simultaneously.
For users, the experience could eventually become much simpler: describe a goal and allow the AI system to manage the underlying workflow.
The technical complexity would increasingly move behind the interface.
Our Perspective
The most interesting part of this development is not simply the headline speed increase.
The larger change is the combination of performance optimisation with multi-agent coordination.
Faster loading makes long-running AI workflows more practical, while automatic delegation could make those workflows more powerful and easier to manage.
However, speed should not be confused with intelligence, and a faster system is not automatically a better system.
The real breakthrough will come if OpenAI can combine speed, intelligent model selection, reliability and low cost into a system that can complete complex work with minimal human supervision.
Conclusion
OpenAI is reportedly introducing major performance improvements across ChatGPT and Codex, including faster handling of extremely long conversations and the expansion of its Multi-Agent V2 architecture.
Reported testing of a 741-round, 231MB conversation showed its average opening time falling from approximately 27.62 seconds to 1.66 seconds, alongside significant reductions in memory use and network requests.
Meanwhile, Multi-Agent V2 is designed to allow a primary agent to divide complex tasks among different AI models and configure their reasoning requirements according to the work involved.
The developments point toward a broader transformation in AI: from individual chat models toward coordinated systems capable of managing complex workflows.
If these improvements scale beyond extreme test cases, the next generation of AI may be defined not only by how intelligent a model is, but by how quickly and efficiently an entire team of AI agents can work together.

