Cloudflare shipped the first two models its Workers AI team has trained in house on October 1, and neither of them writes a word. Clef and Clef-flash take an input state plus a schema of typed questions and hand back a calibrated probability for every permitted answer — weights included, under an Apache 2.0 licence on Hugging Face.
Key takeaways
- Across 10 decision benchmarks a Clef model scores highest on 7, ahead of Typesafe's Jev and other open decision models, according to Cloudflare's own results.
- Over 43 benchmark runs Cloudflare measured Clef at 2.5x Jev's median speed and Clef-flash at 13x, with both models served from GPUs across its edge network.
- Clef is post-trained from a frozen Qwen3.8-27B backbone and Clef-flash from Qwen3.5-9B, both released under Apache 2.0 on Hugging Face.
What you actually send Clef
A request carries a state — a support message, a URL, a user submission — and up to 64 questions, each one of three types. A noul question is a yes/no and comes back as the probability of yes. A choice question picks one option from a set the developer defines and returns the winner, a probability per option and a confidence value. A score question rates the state against an ordered rubric and returns a probability-weighted score plus a probability for each level.
Because the output shape is fixed, there is nothing to parse and no reasoning tokens to wait through. Cloudflare's pitch is that an agent can put this directly in its request path — check "should I take this action?" in tens of milliseconds, then hand off to a language model on Workers AI to actually do the thing.
How Clef differs from Jev
Clef follows the same System One API that Typesafe's Jev uses, so Cloudflare says an existing integration moves over by swapping the endpoint and the model name. Two capability gaps run in Cloudflare's favour: Clef carries a vision encoder and accepts up to four images alongside the state, where Jev is text-only today, and its context window is 64k against Jev's 32k.
The sharpest number Cloudflare offers comes from its own threat intelligence team. Paired with Browser Run, Clef fetched, rendered and classified a website domain in 2.2 seconds; the company's fastest general model in the same workflow, gpt-oss-120b, took 4.7 seconds and returned only two classifications. On Typesafe's own workflow evaluations Cloudflare reports Clef winning three of four areas — invoice processing, customer service and security incidents — which means Jev still takes the fourth. Every figure here is vendor-run, and the launch post publishes them alongside a live demo site rather than third-party verification.
Why the latency gap exists
Clef does not generate tokens one after another. It runs a prefill-only pass over the frozen Qwen backbone, then scores all valid schema choices in parallel through a two-stage attention routing process, deriving answers from internal representations instead of emitting intermediate text. That non-autoregressive decision step is where the speed comes from, not from a smaller model.
Training kept the backbone frozen and jointly optimised a routing head alongside rank-256 LoRA adapters, using label-smoothed cross-entropy for valid outputs and a Brier loss to calibrate the probabilities. Cloudflare also built a secondary objective it calls Reinforcement Learning for Calibrated Decisions, which hands partial credit to adjacent ordinal choices and applies a reference penalty to limit distribution shift. The datasets were synthetic and generated internally, permuting field orders, prompts and schema structures.
The fine-tuning product underneath
The second half of the launch is a reinforcement learning service for tuning Clef to a customer's own workload, beginning as hands-on work with Cloudflare's forward-deployed engineering team and later as a self-serve platform. It stitches together pieces the company already shipped: AI Gateway captures request traffic into a dataset, Workers AI generates rollouts against base Clef, Containers act as the scoring sandbox, a new Trainer updates the weights, and Workers AI's bring-your-own-model path — built on the Cog work that arrived with Cloudflare's Replicate acquisition — redeploys the result.
Cloudflare is its own first customer. It names Trust and Safety submission scoring, support triage, and deciding whether a crawler is a good bot or a bad one, arguing that 15 years of labelled network data is the asset that makes a tuned classifier beat a generic one. The company frames all of it under a mission of being "the agent cloud" — a continuation of the same bet behind its push to rebuild the web for agents. The category itself is barely a month old: Typesafe only exited stealth in September.
FAQ
Is Clef open source?
The weights are. Cloudflare published both Clef and Clef-flash on Hugging Face under an Apache 2.0 licence, which permits commercial use and local deployment. The hosted versions run on Workers AI as @cf/cloudflare/clef and @cf/cloudflare/clef-flash, and Cloudflare says it does not read, store or train on hosted requests unless a customer opts into fine-tuning. The repositories are on Hugging Face as Cloudflare/clef and Cloudflare/clef-flash.
Can an existing Jev integration switch to Clef?
Cloudflare says yes, because Clef implements the same System One API with strictly typed outputs. In practice that means changing the endpoint and the model identifier. Developers reach it through the Workers AI binding or the REST /ai/run path, and AI Gateway sits in front of either.
What is Clef built on?
Both models are post-trained from Alibaba's Qwen family, with the backbone frozen during training: Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash. Cloudflare's earlier public experiment in this area used a different base, adapting DiffusionGemma to expose logprobs as deterministic probabilities, building on independent work by Matt Mastracci in the vLLM project.






