The AI World Keeps Shifting: Modelmaxxing Is Out And Multi-Model Usage Is In
According to a report produced by F5, organizations rely on an average of seven AI models, with 52% of organizations chaining or orchestrating multiple AI models.

Frontier AI models have been all the rage, but many organizations are realizing that less can be more at the moment. In June, Anthropic released Fable 5, one of its most powerful models, but as of August 2026, research indicates it only accounted for 11% of business spending on Anthropic's models, while OpenAI's GPT 5.6 Sol accounted for 23%.
This indicates that the demand for frontier AI models might not warrant the cost for most organizations. After companies like Meta went all-in on tokenmaxxing at the start of the year, many organizations are now modelmaxxing, routing tasks to the most cost-efficient models on offer.
According to a report produced by F5, organizations rely on an average of seven AI models, with 52% of organizations chaining or orchestrating multiple AI models. Similarly, a survey of 1,300 professionals produced by LangChain found model diversity is the norm in the enterprise, with teams routing tasks to different models based on factors like complexity, cost and latency rather than being locked into a single platform.
Under model routing, organizations can evaluate incoming requests and automatically route them to the best model for a given task, making it easier to switch between more expensive proprietary models and lower cost open source models.
"My joke is not everybody needs Fable 5.1 to draft an email, right?," Arsalan Tavakoli, cofounder of Databricks and SVP of Field Engineering told International Business Times in a video interview.
Tavakoli, who says north of 80% of Databrick's customers are multi-model, added its about using the right model for the right task at the right time, something that the startup has aimed to achieve with Unity Gateway, an AI gateway for automatically routing user requests to access models like Claude, GPT, Gemini, Kimi and GLM.
"A very common architecture is either the planner or the advisor is a more powerful model, and you use that to basically build your execution plan, and then you use lower cost models for executing. You don't sacrifice quality, and you actually often times pick up speed at a much lower cost," Tavakoli said.
Tavakoli went on to say that a very common architecture is that either a planner or advisor is a more powerful model, and you can use that to build an execution plan, and then use lower cost models for executing.
"You're now not asking what is the cheapest model, but what is the best model for this task," Jeremy Ung, CTO of Blackline told International Business Times in a video interview. "It's not just the economics of which model is cheapest. It's do the models provide the certainty? Can they run within the harness efficiently? Do they generate the right exhaust that is auditable, and can I trust those outputs."
The desire among enterprise leaders to switch between models is also increasingly driving infrastructure decisions.
"There isn't a one size model approach, and you got to go use a frontier model. It is going to be a real mixture of different types of models, different size models, you know, models that are tuned for specific tasks," Barry Baker, senior vice president of infrastructure at IBM and chief operating officer of IBM infrastructure and general manager of IBM systems, told International Business Times in a video interview.
The need for a mixture of models is moving IBM toward a hybrid architecture, as organizations look to orchestrate and optimize end-to-end GPU, CPU and storage.
"I think we're starting to see the agentic workflow influencing our product roadmaps a lot more," Baker said, noting a need to use models closer to compute to drive a faster
response time and execution for use cases such as upgrading or patching software.
"The last thing I want is that to be an elongated process that's driven by an LLM sitting out in the frontier, right? I want that kind of pushed onto the platform so I can handle that and manage it in real time," Baker concluded.
© Copyright IBTimes 2026. All rights reserved.

