AI/News
Cloudflare releases Clef, two open-source "decision models" that return typed answers with probabilities in 39 to 209 milliseconds
Clef and Clef-flash run on Workers AI and sit on Hugging Face under Apache 2.0; Cloudflare says Clef leads the Jev Decision Index and classified a website in 2.2 seconds against 4.7 for its fastest general model. A fine-tuning service comes with them.
By Daily Aletheia · Checked against the primary source · 2 October 2026 · 2 min read

Cloudflare announced Clef and Clef-flash on 1 October, in a post by Michelle Chen, Alex Reneau and Kevin Flansburg. A decision model, in Cloudflare's description, "makes classifications to help agents decide how to act, based on certain probabilities": pass it a support message and ask whether it is urgent and which team should take it, and it returns typed answers with probabilities that code can act on - route, escalate, or hand to a human.
What is released
Both models are hosted on Workers AI and open-sourced on Hugging Face under an Apache 2.0 licence. Clef is built on a frozen Qwen3.8-27B and Clef-flash on Qwen3.5-9B, each with a routing head and low-rank adapters trained on Cloudflare's own synthetic datasets. Unlike a chat model they generate no text: a single pass scores the valid answers in parallel, which is where the speed comes from. Clef has a vision encoder, so it can classify images, and a 64,000-token context window.
Cloudflare's numbers
Across the benchmarks Cloudflare chose from the Jev Decision Index, Clef scores 98.47 on BFCL and 94.20 on BANKING77 against Jev's 95.75 and 79.74; on When2Call Jev leads, 80.97 to 72.37. Median latency is 209.3 milliseconds for Clef and 38.8 for Clef-flash, against 524.1 for Jev. On Typesafe's own four-workflow suite Cloudflare says Clef beats Jev in three of four. Its Threat Intelligence team used Clef to fetch, render and classify a website domain in 2.2 seconds, where gpt-oss-120b took 4.7 seconds in the same workflow.
The fine-tuning service
Cloudflare is also offering to fine-tune Clef on a customer's own data - first through its forward-deployed engineering team, later as a self-serve platform built from AI Gateway to capture requests, Workers AI for rollouts, Containers as the sandbox and a new Trainer to update the weights. The company says it does not read, store or train on requests to the hosted models unless a customer opts into fine-tuning.
Every figure here is Cloudflare's own, from its post and its live benchmark site; the comparisons were run by the company releasing the model.