Cloudflare has quietly integrated Kimi K2.5 from Moonshot AI into its Workers AI platform, setting it as the default model for the Agents SDK starter. According to the official blog, the model is already handling real security audit tasks internally — processing over 7 billion tokens per day and slashing costs by 77% compared to mid-tier commercial models, saving nearly $1.85 million annually.
Security Audit Agent: 7 Billion Tokens Per Day
Cloudflare engineers deployed Kimi K2.5 as the primary model for programming agents in the OpenCode environment, along with a public code review agent named Bonk connected to automated pipelines. But the most impressive numbers come from internal security audits: the agent handles over 7 billion tokens daily. Running the same workload on a standard commercial model would cost about $2.4 million per year; switching to Kimi K2.5 cuts that by 77%, saving $1.85 million. These figures come straight from Cloudflare's official blog.
Kimi K2.5 is one of the few open-source models offering frontier specs: 256K context window, multi-turn tool calling, vision input, and structured output — making it highly practical for long-context reasoning agent tasks.
Three Platform-Level Improvements for Long Conversations
Beyond the model swap, Cloudflare rolled out three optimizations targeting cost and efficiency in long agent conversations:
- Prefix Caching discounts: Already-processed input tokens in multi-turn dialogs are not billed again; cached tokens get a discount, saving significantly on long-running tasks.
- Session Affinity Header: A new
x-session-affinityrequest header routes the same session to the same model, boosting cache hit rates. OpenCode and the Agents SDK starter already support it. - Async batch inference API: Requests exceeding sync rate limits can be queued asynchronously, typically completing within 5 minutes in internal tests. Ideal for code scanning and research tasks that don't need real-time responses.
Under the Hood: Custom Infire Inference Engine
Rather than using off-the-shelf inference frameworks, Cloudflare built a custom core with its own Infire engine, leveraging data parallelism, tensor parallelism, and expert parallelism, plus a separated prefix processing architecture. Kimi K2.5 is the first large model inference deployment on Workers AI, signaling Cloudflare's ambition in AI infrastructure — tightly integrated with its network platform and aggressively priced.

