Coinbase Cuts AI Spending by Nearly 50%: Default Open-Weight Models, Smart Routing, and Caching Strategies

Coinbase Cuts AI Spending by Nearly 50%: Default Open-Weight Models, Smart Routing, and Caching Strategies

N
News Editor
2026-06-27 19:36:35
Coinbase CEO Brian Armstrong disclosed that the company has slashed AI spending by nearly half through a series of infrastructure optimizations, while token usage continues to grow. Key measures include defaulting to open-weight models like GLM 5.2 and Kimi 2.7, implementing an LLM gateway for intelligent model routing based on task requirements, and deploying cache-aware architecture that boosted LibreChat's cache hit rate from 5% to 60%. The company also requires engineers to keep context lean to reduce inference costs. Armstrong emphasized that the goal is not to cap AI usage but to build infrastructure able to support exponential growth, and that model selection should eventually be automated by AI itself.
CoinbaseAI cost controlopen-weight modelsLLM gatewaymodel routingcache optimizationBrian Armstrongartificial intelligence

Background: A New Approach to Soaring AI Costs

On June 27, Coinbase CEO Brian Armstrong took to social media to share the company's breakthrough practices in AI cost management. As token usage grows exponentially, AI spending pressures would typically rise in tandem, but Coinbase chose not to rely on traditional usage caps or spending alerts. Instead, the company focused on infrastructure-level improvements. Armstrong stressed that the key lies not in adding friction, but in building better default models, routing mechanisms, and caching systems.

Three Core Strategies: Default Models, Smart Routing, and Caching

First, Coinbase has switched its default models to open-weight alternatives such as GLM 5.2 and Kimi 2.7 through an LLM gateway, while still encouraging engineers to choose the most suitable model for each task. Data shows that 91% of employees had never hit the usage limit, so instead of lowering quotas or adding reminders, the company simply shifted to lower-cost defaults, significantly reducing baseline inference costs.

Second, on model routing, Coinbase pre-processes prompts in a custom pipeline and routes tasks to the most cost-effective model based on cache hit rates and model pricing. For example, the planning phase may require a frontier model, but using one during execution is often overkill. Armstrong believes that in the future, humans should not manually select models—AI should handle that automatically, enabling finer-grained resource allocation.

Third, caching is what Armstrong calls 'the easiest way to drive up costs.' All Coinbase requests are cache-aware, maximizing reuse of hot cache. For instance, after properly implementing caching for LibreChat, the hit rate soared from 5% to 60%, dramatically reducing redundant computation. Additionally, Coinbase requires engineers to keep context lean: start new sessions when switching tasks, narrow file context scopes, and disconnect unused tools.

Results and Outlook: Spending Halved While Growth Continues

Through these practices, Coinbase has cut its AI spending by nearly half while token usage continues to rise. Armstrong emphasized that the goal is not to suppress AI usage, but to build infrastructure capable of supporting exponential growth. This case provides a replicable cost optimization paradigm for the crypto industry and tech companies: during the AI application boom, intelligent architectural management can achieve a virtuous cycle of 'more usage for less cost.'

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
300

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.