AI/News

Cloudflare releases Clef, two open-source "decision models" that return typed answers with probabilities in 39 to 209 milliseconds

Clef and Clef-flash run on Workers AI and sit on Hugging Face under Apache 2.0; Cloudflare says Clef leads the Jev Decision Index and classified a website in 2.2 seconds against 4.7 for its fastest general model. A fine-tuning service comes with them.

By Daily Aletheia · Checked against the primary source · 2 October 2026 · 2 min read


Image: Cloudflare - the illustration from "Introducing Clef: our open-source decision models, and new RL fine-tuning platform", blog.cloudflare.com, 1 October 2026, cropped to the artwork

Cloudflare announced Clef and Clef-flash on 1 October, in a post by Michelle Chen, Alex Reneau and Kevin Flansburg. A decision model, in Cloudflare's description, "makes classifications to help agents decide how to act, based on certain probabilities": pass it a support message and ask whether it is urgent and which team should take it, and it returns typed answers with probabilities that code can act on - route, escalate, or hand to a human.

What is released

Both models are hosted on Workers AI and open-sourced on Hugging Face under an Apache 2.0 licence. Clef is built on a frozen Qwen3.8-27B and Clef-flash on Qwen3.5-9B, each with a routing head and low-rank adapters trained on Cloudflare's own synthetic datasets. Unlike a chat model they generate no text: a single pass scores the valid answers in parallel, which is where the speed comes from. Clef has a vision encoder, so it can classify images, and a 64,000-token context window.

Cloudflare's numbers

Across the benchmarks Cloudflare chose from the Jev Decision Index, Clef scores 98.47 on BFCL and 94.20 on BANKING77 against Jev's 95.75 and 79.74; on When2Call Jev leads, 80.97 to 72.37. Median latency is 209.3 milliseconds for Clef and 38.8 for Clef-flash, against 524.1 for Jev. On Typesafe's own four-workflow suite Cloudflare says Clef beats Jev in three of four. Its Threat Intelligence team used Clef to fetch, render and classify a website domain in 2.2 seconds, where gpt-oss-120b took 4.7 seconds in the same workflow.

The fine-tuning service

Cloudflare is also offering to fine-tune Clef on a customer's own data - first through its forward-deployed engineering team, later as a self-serve platform built from AI Gateway to capture requests, Workers AI for rollouts, Containers as the sandbox and a new Trainer to update the weights. The company says it does not read, store or train on requests to the hosted models unless a customer opts into fine-tuning.

Every figure here is Cloudflare's own, from its post and its live benchmark site; the comparisons were run by the company releasing the model.

Sources & further reading

  1. 01Cloudflare - Introducing Clef: our open-source decision models, and new RL fine-tuning platform, 1 Oct 2026 ↗Primary