Cloudflare has released Clef and Clef-flash, a pair of decision models built on Alibaba’s open-source Qwen models, 16 days after TypeSafe AI introduced Jev. The company says the new models are API-compatible with Jev, offer open weights, and can process images. In Cloudflare’s own published tests, Clef leads on most benchmarks. Public pricing on Cloudflare’s side, though, shows Clef at roughly 5.7 times Jev’s listed unit cost.
Cloudflare follows Jev with its own decision model lineup
TypeSafe AI launched Jev on Sept. 15 and framed “decision models” as a standalone product category. On Oct. 1, Cloudflare answered with Clef, using Qwen as the foundation. The report notes that Jev was previously examined by BlockTempo, and that founder Diogo Almeida is a former OpenAI researcher. Jev is described by its creator as a System One Model.
A decision model, in simple terms, does not generate a paragraph of text. Instead, it takes an input and a fixed set of valid choices, then returns a probability for each option. The example in the report is customer support triage: the system asks whether a message is urgent and which team should handle it, then routes the ticket by probability or hands it to a human when confidence is low. Where general-purpose large language models produce open-ended outputs that may vary from run to run, decision models are built for bounded outputs, speed, and low cost.
Clef and Clef-flash are on Workers AI with Apache 2.0 weights
Clef and the smaller Clef-flash, both launched on Oct. 1, are the first machine learning models trained in-house by the Cloudflare Workers AI team. Both are now available through Workers AI. Their weights are licensed under Apache 2.0 and published on Hugging Face for download.
Cloudflare says the API is fully compatible with Jev. The company also says it will not read, retain, or use customer requests and responses for training, except when a customer chooses to use fine-tuning services. The name Clef comes from the musical symbol that defines pitch at the beginning of a staff, and the initials “CF” also echo Cloudflare’s name.
Built on Qwen, with vision support and 64K context
Clef is based on Qwen3.8-27B and is listed on its Hugging Face model card at about 27.36 billion parameters. Clef-flash is based on Qwen3.5-9B, at about 9.41 billion. Cloudflare says it freezes the base model, trains only a routing head, and adds rank-256 LoRA layers. In practical terms, the base stays intact while a relatively small adjustment layer is added on top.
At inference time, the base model reads the input once and scores predefined questions and answer choices in parallel under a schema, rather than generating output token by token. That means there is no intermediate text generation step.
Both Clef and Clef-flash include a vision encoder and can read images. Their context window is 64K. Cloudflare says Jev currently handles text only and has a 32K context window.
Cloudflare’s published benchmarks show Clef leading most quality tests
The report states that all of the following results were run and published by Cloudflare itself. Cloudflare says Clef is currently ahead on the “Jev Decision Index.” In the 10 quality evaluations listed in the company’s blog post, Clef or Clef-flash posted the top result in 7. Jev led in 2. In the remaining phishing-detection benchmark, PhishNChips, a model adapted from Google DiffusionGemma scored 85.35, ahead of Clef’s 79.60.
Some of the margins were large. On BANKING77, a bank customer-service intent classification task, Clef scored 94.20 while Jev scored 79.74. On Home appliances, Clef-flash scored 97.73 and Jev scored 52.27.
Jev kept the lead on two tasks. On When2Call, Jev scored 80.97 versus Clef’s 72.37. On BRIGHT, Jev scored 47.52 and Clef scored 45.91.
Clef-flash also showed an obvious weak spot. On CLINC150+OOS, it scored just 66.77, compared with 89.27 for Jev and 97.43 for Clef.
Cloudflare also cited TypeSafe’s own workflow benchmark set and said Clef won 3 of 4 tasks, though by narrow margins: invoice processing at 64.7 versus 61.8, customer service at 76.3 versus 76.0, and security incidents at 62.9 versus 61.7. Jev led on agent trajectory observability, 71.6 to Clef’s 68.5.
Latency is where Clef opens the gap
Across 43 evaluations, median latency was 209.3 milliseconds for Clef, 38.8 milliseconds for Clef-flash, and 524.1 milliseconds for Jev. On p95, meaning 95% of requests complete within that time, the figures were 238.6, 122.4, and 536.0 milliseconds, respectively.
The open-source model Laya posted a median latency of just 5.8 milliseconds, but ranked last on quality. Cloudflare explicitly noted that the speed came at the expense of output quality.
The company’s internal threat-intelligence example was more direct. Using Clef with Browser Run for domain classification, the combined fetch-render-classify workflow took 2.2 seconds. The same flow, when switched to its fastest general-purpose LLM, gpt-oss-120b, took 4.7 seconds and returned only two classifications.
Listed pricing puts Clef at about 5.7x Jev’s unit cost
Cloudflare’s blog post did not spell out pricing in the body text, but the Workers AI pricing page lists Clef at $0.24 per million input tokens and Clef-flash at $0.09. TypeSafe lists Jev at $0.042 per million input tokens, with output tokens free.
On that basis, Clef’s listed unit price is about 5.7 times Jev’s, while Clef-flash comes in at around 2.1 times Jev’s.
The report also says the comparison should not be reduced to unit pricing alone. Clef’s weights can be downloaded for free and self-hosted. TypeSafe has not made Jev’s weights public.
Cloudflare is pairing the model launch with RL fine-tuning services
According to the report, Cloudflare is aiming to sell more than the model itself. The company launched a reinforcement learning, or RL, fine-tuning service alongside Clef. At first, customers will be guided by the Forward Deployed Engineer, or FDE, team. Cloudflare then plans to turn the process into a self-serve offering so customers can collect data, fine-tune, and redeploy on their own.
The workflow is described as tying into Cloudflare’s existing products. AI Gateway turns passing AI requests into datasets, while Workers AI generates rollouts and connects to later deployment steps. The source text is cut off in this section, and no further details are provided.

