Coinbase Cuts AI Spending by Half via LLM Gateway, Default Open-Weight Models, and Smart Caching

Coinbase Cuts AI Spending by Half via LLM Gateway, Default Open-Weight Models, and Smart Caching

N
News Editor
2026-06-27 19:36:35
Coinbase CEO Brian Armstrong披露公司通过默认使用GLM 5.2、Kimi 2.7等开放权重模型、优化模型路由与缓存策略,将AI支出削减近50%,同时Token使用量仍保持增长。91%员工未触及使用上限,因此公司转向低成本默认模型。缓存命中率从5%提升至60%,并强调未来应由AI自动选择模型。
CoinbaseAI spendingmodel routingcaching optimizationGLM 5.2Kimi 2.7LLM gatewaycost reduction

Coinbase Halves AI Costs: The Key Is Not Usage Friction but Infrastructure Optimization

On June 27, Coinbase CEO Brian Armstrong shared the company's core strategy for controlling AI spending. He argued that to keep AI costs stable while usage grows exponentially, the solution is not to introduce usage friction or spending alerts, but to optimize the underlying infrastructure: better default models, intelligent routing, and caching mechanisms. Through its self-developed LLM Gateway, Coinbase has reduced AI spending by nearly 50%, while token consumption continues to rise.

Default Model Migration and Routing: GLM 5.2 and Kimi 2.7 Take Center Stage

Coinbase's LLM Gateway now defaults to open-weight models such as GLM 5.2 and Kimi 2.7, while still allowing engineers to choose the best model per task. Data shows that 91% of employees never hit their usage limits, so instead of lowering quotas or adding warnings, the company shifted to cheaper default models. For routing, Coinbase preprocesses prompts in custom flows and routes tasks to the most cost-effective model based on cache hit rates and model pricing. For example, planning stages may require frontier models, but execution stages often do not. Armstrong believes humans should not manually select models in the future; AI can handle that automatically.

Cache Strategy and Context Pruning: Key Details for Cost Control

Armstrong highlighted that cache misses are the easiest way to drive up AI costs. All Coinbase requests are cache-aware to maximize reuse of hot caches. For instance, after implementing proper caching for LibreChat, the cache hit rate jumped from 5% to 60%. Additionally, Coinbase requires engineers to keep context lean, including starting new sessions when switching tasks, narrowing file context scopes, and disconnecting unused tools. The goal is not to suppress AI usage but to build infrastructure that can support exponential growth. Through these measures, Coinbase has cut AI spending by nearly half while token usage continues to grow, validating the effectiveness of the approach.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
300

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.