
You’re reading an issue of “The AI Economy,” my newsletter exploring the forces shaping the AI era—tracking how AI is rewriting business, work, technology, and culture. Subscribe to get expert insights and curated updates delivered straight to your inbox.
Snowflake has announced a feature that will soon automatically select the best AI model to handle a customer’s request—and it may not even be one from Anthropic or OpenAI. Dynamic model routing, in Cortex AI as a private preview, is the AI data cloud company’s answer to enterprises that have moved away from the one-model-fits-all approach, routing simpler tasks to cheaper models and reserving frontier models for work that needs them.
Along with dynamic model routing, Snowflake’s catalog of open models will be expanding. The company said it will add support for DeepSeek V4 Flash 0731 and Z.ai’s GLM-5.3.
Entering the ‘Intelligence Efficiency’ Era
At the center of today’s announcement is Snowflake’s Cortex AI Gateway, the control layer governing how AI agents access data, tools, and models. Powered by dynamic model routing, Cortex AI evaluates tasks and automatically selects the best-suited model based on quality and cost. Repetitive and less complex tasks are directed to more efficient models, while assignments requiring deeper reasoning access frontier models. Snowflake said dynamic model routing is integrated with its “flagship AI” products, including CoCo and CoWork. It’s also available to Cortex AI-connected third-party agents.
While there may have been an initial push to increase AI adoption in the workplace, companies are now taking a more measured approach. We’ve moved past the “tokenmaxxing” phase and are now into what others, such as IBM, have coined “valuemaxxing.” Leaders want to use AI, but don’t want to break the bank doing so. Dynamic model routing promises a more economically friendly option: Have AI determine the best model.
“Companies are investing aggressively in AI, but they are also becoming more rigorous about the economics,” said Sridhar Ramaswamy, Snowflake’s chief executive, in a blog post. “It is no longer enough to show that employees are using AI, or that token consumption is growing. Usage is an input. The question that matters is what a company gets in return.”
He said we’re in an era he terms “intelligence efficiency,” where companies want to move fast, save on cost, and deliver real, measurable business value. “Achieving this efficiency requires model flexibility, both in the freedom to use the best models for the job and in the ability to leverage intelligent routing to match each task to the right model based on cost, speed, and performance,” Ramaswamy said.
This presents a “huge unlock for enterprises,” he continued. “As the cost per task falls, the threshold for where it makes sense to apply AI falls with it.” Ramaswamy reasoned that because workflows are now more economically viable, companies have greater flexibility to experiment with and deploy AI across more parts of their organizations.
Underscoring his point, Snowflake shared internal testing it said demonstrated how mixing open and proprietary models can result in comparable quality and improved token efficiency. When agents built a dbt pipeline, the company claimed that dynamic model routing achieved up to three times greater token efficiency and the same quality as using only a frontier model. In another test, Snowflake reported that its engineering team was able to make the same number of pull requests with 25 percent higher token efficiency.
But Snowflake isn’t the only one that has shipped a model routing feature. It joins a growing list of tech providers offering this service, including Amazon Web Services, Databricks, Microsoft, and OpenRouter.
Chinese AI Models Come to Snowflake
In addition to dynamic model routing, Snowflake’s portfolio of third-party AI models continues to grow. Along with those from Anthropic, OpenAI, Google, SpaceXAI, Meta, and Mistral, the company is adding two Chinese models: DeepSeek V4 Flash 0731 and GLM-5.3. Customers can use them across Cortex AI, Snowflake CoCo, and Snowflake CoWork.
Notably, DeepSeek V4 Flash is available in private preview, while GLM-5.3 is coming soon in private preview.
DeepSeek V4 Flash was first announced in April and moved into public beta in July. It’s one of two variations of the Chinese AI lab’s flagship model, with 284 billion parameters, and is designed for volume and speed. Leaderboards have it as a top performer, though one evaluation found it struggled in real-world testing tasks. Before August 16, it cost $0.14 per million tokens input and $0.28 per million tokens output—making it more affordable than GPT-5.4 Nano and Gemini 3.1 Flash-Lite. However, DeepSeek has since raised prices, with V4 Flash now costing $1.32 per million output tokens during peak times.
GLM-5.3 comes from Z.ai and was released last week. It generated headlines for its improved long-horizon coding and cybersecurity capabilities, putting it in closer competition with OpenAI and Anthropic. It has the same base model as its predecessor, GLM-5.2, and has 743 billion parameters. Based on Z.ai’s testing, GLM-5.3 outperformed Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol on the CyberGym benchmark testing vulnerability identification and validation. However, it lagged in ExploitBench—scoring 54.4 percent compared with 78 percent for Mythos 5 and 76.5 percent for GPT-5.6 Sol.
Snowflake’s own team evaluated both models and found that DeepSeek V4 Flash outperformed leading proprietary models in its testing on enterprise-focused tasks, scoring 74.4 percent on data engineering tasks. And though it’s adding GLM-5.3, the company evaluated GLM-5.2 and reported that it “performed strongly,” with a score of 62.8 percent. It also noted that GLM-5.2 consumed fewer tokens than any other model tested.
For Ramaswamy, “this freedom and flexibility are fundamental to how we think about the Agentic Enterprise. As models continue to evolve, enterprises should be able to benefit from that innovation across the full scope of their data and context.”
