Cloudflare said on Oct. 1 that it released two internally trained AI decision models, Clef and Clef-flash. Both are now available on Workers AI, and their weights have been open-sourced on Hugging Face under the Apache 2.0 license. The company also introduced a reinforcement learning service for enterprises that want to fine-tune Clef for their own workloads.
Built for classification rather than text generation
Cloudflare said decision models are designed to classify inputs and help AI agents choose the next action, instead of generating text. In one example, a customer support message can be passed to the model to determine whether it is urgent and which team should handle it. The model then returns a typed answer with probabilities, allowing software to route the ticket, escalate it, or send it to a human operator.
Cloudflare said discussion around this category of models has picked up in recent weeks following the release of Jev by Typesafe AI.
Performance figures shared by Cloudflare
According to Cloudflare, its threat intelligence team has already used Clef to classify website domains. In one example, the model judged that a domain had a 95% probability of being a fashion site, an 85% probability of being an e-commerce site, and less than a 1% probability of being a phishing site.
Cloudflare said Clef completed crawling, rendering, and classification in 2.2 seconds. The same workflow took 4.7 seconds with the company’s fastest general-purpose model, gpt-oss-120b, and returned only two classifications.
In tests published by Cloudflare, Clef led the Jev Decision Index. Median latency came in at 209.3 milliseconds for Clef, 38.8 milliseconds for Clef-flash, and 524.1 milliseconds for Jev. Cloudflare also said Clef adds a vision encoder compared with Jev, allowing it to classify image content, extends context length from 32k to 64k, and remains fully compatible with the Jev API.
Model backbone and inference design
On the technical side, Clef is built on a frozen Qwen3.8-27B backbone, while Clef-flash uses Qwen3.5-9B, followed by post-training. Cloudflare said inference runs with a single prefill step and then scores all valid options in parallel. Because the system does not need to generate intermediate text token by token, it runs much faster than a standard autoregressive language model.
Reinforcement learning service starts with engineering support
Cloudflare said its new reinforcement learning service will first be delivered with help from its deployment engineering team, which will work with customers to fine-tune Clef for specific workloads. The company said it plans to turn that into a self-serve platform later, allowing customers to collect data, fine-tune models, and redeploy them on their own.
Internally, Cloudflare said teams are already interested in using the model to review trust and safety reports, route customer support requests, and judge whether a crawler is benign or malicious.
Developer reaction on X
Developer Peter Steinberger reposted the announcement on X and wrote, “I’ve never seen an idea spread this fast.”

