
How do U.S. tech companies control AI costs amid exponential token consumption growth in 2026? Coinbase has delivered a surprising answer—making Chinese open-source models the enterprise default.
On June 27, Coinbase CEO Brian Armstrong disclosed on X that the company has cut AI spending by nearly 50% by making open-weight models—including Zhipu AI's GLM 5.2 and Moonshot AI's Kimi 2.7—the default options, combined with smart routing and aggressive caching strategies. Token usage continues to grow.
91% of Employees Never Hit Quotas—the Problem Isn't Controls, It's Cost
Armstrong revealed that internal data shows 91% of Coinbase engineers have never reached their AI usage quotas. Rather than tightening controls or sending frequent budget reminders, the company shifted to "cheaper default models".
"The key to cost control is not limiting usage, but optimizing default model selection, task routing mechanisms, and caching strategies," Armstrong wrote.
Three Cost-Saving Measures: Smart Routing, Aggressive Caching, Context Discipline
Coinbase implemented a systematic cost-optimization solution through its internal LLM gateway:
1. Intelligent Model Routing
The system preprocesses each prompt and automatically routes tasks to the most cost-effective model based on cache hit probability and per-token pricing across different models. Armstrong believes that while complex tasks like planning and reasoning may require frontier models, execution-level tasks don't necessarily need higher-cost models. In the future, model selection should be automated by AI rather than left to manual decisions.
2. Aggressive Caching Strategy
All AI requests at Coinbase are now cache-aware—checking whether previous responses can be reused before generating new ones. After optimization, LibreChat's cache hit rate jumped from 5% to 60%.
3. Context Discipline
Engineers are required to keep context windows lean—starting new sessions, narrowing file scope, disconnecting unused tools—to reduce wasted token consumption.
Open-Weight Models as Defaults: Chinese AI's "Commercial Validation" in Silicon Valley
Making GLM 5.2 and Kimi 2.7 the default enterprise models marks a milestone for Chinese open-source AI in Western enterprise infrastructure.
The business logic is clear and compelling: cutting AI spending by nearly 50% while maintaining exponential token consumption growth means Coinbase has, to a significant degree, decoupled consumption growth from cost increases.
For U.S. AI providers like OpenAI and Anthropic, this constitutes a real challenge—when enterprises discover that everyday tasks can be handled by Chinese open-source models at extremely low costs, the "premium pricing" for frontier models is being squeezed.
Not an "All-or-Nothing" Approach: Frontier Models Still Used for Complex Tasks
It's worth noting that Coinbase hasn't completely abandoned U.S. frontier models. Engineers can still call frontier models for complex planning and reasoning tasks, while code review employs a multi-model parallel strategy where outputs cross-validate each other.
This "open-source by default + on-demand premium" hybrid model may become the new paradigm for enterprise AI deployment—routine tasks handled by low-cost models, high-stakes tasks reserved for top-tier models, driving down total costs without sacrificing performance on critical applications.
Editor's Note
The Coinbase case is real-world validation of the "Chinese AI cost-performance" approach in the global market. When U.S. frontier model API prices remain high while Chinese open-source models continue climbing international leaderboards, economic rationality is overtaking tech politics.
Armstrong's post made no mention of the "geopolitical attributes" of model sources—only cost and efficiency. This suggests that in global enterprise AI procurement decisions, "how well it works and how much it costs" is regaining its position as the decisive factor.
From Silicon Valley to Europe, from Coinbase to Pinterest, a growing number of overseas enterprises are "voting with their feet" to redefine the competitive rules of the global AI industry. As cumulative global downloads of Chinese open-source models like Qwen and DeepSeek surpass 10 billion, this trend is moving from exception to norm.
This article is based on reports from Coinbase CEO Brian Armstrong's X post, ChainCatcher, Edgen.Tech, Gate.com, and other media outlets.