Cloudflare Integrates Kimi K2.5: Security Audit Agent Processes 7 Billion Tokens Daily, Slashes Costs by 77%

Cloudflare Integrates Kimi K2.5: Security Audit Agent Processes 7 Billion Tokens Daily, Slashes Costs by 77%

N
News Editor 01
2026-07-24 07:50:15
Cloudflare's Workers AI platform integrates Moonshot AI's Kimi K2.5, powering an internal security audit agent that handles over 7 billion tokens daily, cutting costs by 77% vs. mid-tier commercial models.

Cloudflare has quietly integrated Kimi K2.5 from Moonshot AI into its Workers AI platform, setting it as the default model for the Agents SDK starter. According to the official blog, the model is already handling real security audit tasks internally — processing over 7 billion tokens per day and slashing costs by 77% compared to mid-tier commercial models, saving nearly $1.85 million annually.

Security Audit Agent: 7 Billion Tokens Per Day

Cloudflare engineers deployed Kimi K2.5 as the primary model for programming agents in the OpenCode environment, along with a public code review agent named Bonk connected to automated pipelines. But the most impressive numbers come from internal security audits: the agent handles over 7 billion tokens daily. Running the same workload on a standard commercial model would cost about $2.4 million per year; switching to Kimi K2.5 cuts that by 77%, saving $1.85 million. These figures come straight from Cloudflare's official blog.

Kimi K2.5 is one of the few open-source models offering frontier specs: 256K context window, multi-turn tool calling, vision input, and structured output — making it highly practical for long-context reasoning agent tasks.

Three Platform-Level Improvements for Long Conversations

Beyond the model swap, Cloudflare rolled out three optimizations targeting cost and efficiency in long agent conversations:

  • Prefix Caching discounts: Already-processed input tokens in multi-turn dialogs are not billed again; cached tokens get a discount, saving significantly on long-running tasks.
  • Session Affinity Header: A new x-session-affinity request header routes the same session to the same model, boosting cache hit rates. OpenCode and the Agents SDK starter already support it.
  • Async batch inference API: Requests exceeding sync rate limits can be queued asynchronously, typically completing within 5 minutes in internal tests. Ideal for code scanning and research tasks that don't need real-time responses.

Under the Hood: Custom Infire Inference Engine

Rather than using off-the-shelf inference frameworks, Cloudflare built a custom core with its own Infire engine, leveraging data parallelism, tensor parallelism, and expert parallelism, plus a separated prefix processing architecture. Kimi K2.5 is the first large model inference deployment on Workers AI, signaling Cloudflare's ambition in AI infrastructure — tightly integrated with its network platform and aggressively priced.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
1800

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.