AI Newsway

Cloudflare's First In-House Model Leads 7 of 10 Decision Benchmarks — and Ships Open

Clef and Clef-flash return calibrated probabilities instead of text, run on Workers AI's edge GPUs, and arrive with a reinforcement learning tuning service.

|5 min read0
AI Summary
Cloudflare released Clef and Clef-flash on October 1, the first models trained by its Workers AI team. They return calibrated probabilities for typed questions rather than generating text, lead 7 of 10 decision benchmarks in Cloudflare's own testing, and ship under Apache 2.0 on Hugging Face from frozen Qwen backbones. A reinforcement learning service for tuning Clef to customer workloads launches alongside them, beginning as hands-on engineering work.
Cloudflare's San Francisco office, where the Workers AI team released Clef and Clef-flash, the company's first internally trained models, on October 1.
Cloudflare's San Francisco office, where the Workers AI team released Clef and Clef-flash, the company's first internally trained models, on October 1.

Cloudflare shipped the first two models its Workers AI team has trained in house on October 1, and neither of them writes a word. Clef and Clef-flash take an input state plus a schema of typed questions and hand back a calibrated probability for every permitted answer — weights included, under an Apache 2.0 licence on Hugging Face.

Key takeaways

  • Across 10 decision benchmarks a Clef model scores highest on 7, ahead of Typesafe's Jev and other open decision models, according to Cloudflare's own results.
  • Over 43 benchmark runs Cloudflare measured Clef at 2.5x Jev's median speed and Clef-flash at 13x, with both models served from GPUs across its edge network.
  • Clef is post-trained from a frozen Qwen3.8-27B backbone and Clef-flash from Qwen3.5-9B, both released under Apache 2.0 on Hugging Face.

What you actually send Clef

A request carries a state — a support message, a URL, a user submission — and up to 64 questions, each one of three types. A noul question is a yes/no and comes back as the probability of yes. A choice question picks one option from a set the developer defines and returns the winner, a probability per option and a confidence value. A score question rates the state against an ordered rubric and returns a probability-weighted score plus a probability for each level.

Because the output shape is fixed, there is nothing to parse and no reasoning tokens to wait through. Cloudflare's pitch is that an agent can put this directly in its request path — check "should I take this action?" in tens of milliseconds, then hand off to a language model on Workers AI to actually do the thing.

How Clef differs from Jev

Clef follows the same System One API that Typesafe's Jev uses, so Cloudflare says an existing integration moves over by swapping the endpoint and the model name. Two capability gaps run in Cloudflare's favour: Clef carries a vision encoder and accepts up to four images alongside the state, where Jev is text-only today, and its context window is 64k against Jev's 32k.

The sharpest number Cloudflare offers comes from its own threat intelligence team. Paired with Browser Run, Clef fetched, rendered and classified a website domain in 2.2 seconds; the company's fastest general model in the same workflow, gpt-oss-120b, took 4.7 seconds and returned only two classifications. On Typesafe's own workflow evaluations Cloudflare reports Clef winning three of four areas — invoice processing, customer service and security incidents — which means Jev still takes the fourth. Every figure here is vendor-run, and the launch post publishes them alongside a live demo site rather than third-party verification.

Why the latency gap exists

Clef does not generate tokens one after another. It runs a prefill-only pass over the frozen Qwen backbone, then scores all valid schema choices in parallel through a two-stage attention routing process, deriving answers from internal representations instead of emitting intermediate text. That non-autoregressive decision step is where the speed comes from, not from a smaller model.

Training kept the backbone frozen and jointly optimised a routing head alongside rank-256 LoRA adapters, using label-smoothed cross-entropy for valid outputs and a Brier loss to calibrate the probabilities. Cloudflare also built a secondary objective it calls Reinforcement Learning for Calibrated Decisions, which hands partial credit to adjacent ordinal choices and applies a reference penalty to limit distribution shift. The datasets were synthetic and generated internally, permuting field orders, prompts and schema structures.

The fine-tuning product underneath

The second half of the launch is a reinforcement learning service for tuning Clef to a customer's own workload, beginning as hands-on work with Cloudflare's forward-deployed engineering team and later as a self-serve platform. It stitches together pieces the company already shipped: AI Gateway captures request traffic into a dataset, Workers AI generates rollouts against base Clef, Containers act as the scoring sandbox, a new Trainer updates the weights, and Workers AI's bring-your-own-model path — built on the Cog work that arrived with Cloudflare's Replicate acquisition — redeploys the result.

Cloudflare is its own first customer. It names Trust and Safety submission scoring, support triage, and deciding whether a crawler is a good bot or a bad one, arguing that 15 years of labelled network data is the asset that makes a tuned classifier beat a generic one. The company frames all of it under a mission of being "the agent cloud" — a continuation of the same bet behind its push to rebuild the web for agents. The category itself is barely a month old: Typesafe only exited stealth in September.

FAQ

Is Clef open source?

The weights are. Cloudflare published both Clef and Clef-flash on Hugging Face under an Apache 2.0 licence, which permits commercial use and local deployment. The hosted versions run on Workers AI as @cf/cloudflare/clef and @cf/cloudflare/clef-flash, and Cloudflare says it does not read, store or train on hosted requests unless a customer opts into fine-tuning. The repositories are on Hugging Face as Cloudflare/clef and Cloudflare/clef-flash.

Can an existing Jev integration switch to Clef?

Cloudflare says yes, because Clef implements the same System One API with strictly typed outputs. In practice that means changing the endpoint and the model identifier. Developers reach it through the Workers AI binding or the REST /ai/run path, and AI Gateway sits in front of either.

What is Clef built on?

Both models are post-trained from Alibaba's Qwen family, with the backbone frozen during training: Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash. Cloudflare's earlier public experiment in this area used a different base, adapting DiffusionGemma to expose logprobs as deterministic probabilities, building on independent work by Matt Mastracci in the vLLM project.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

Four Opus 5 Request Patterns Now Return 400 on Claude Opus 5.5
Developer Tools

Four Opus 5 Request Patterns Now Return 400 on Claude Opus 5.5

Claude Opus 5.5 is 20% cheaper than Opus 5 and returns HTTP 400 on four request patterns Opus 5 accepted. A fifth change silences agent progress streams.

Seung Jung4 days ago
Docker Moves Its AI Agent Sandboxes to the Cloud, Billed by the Second From $0.07 an Hour
Developer Tools

Docker Moves Its AI Agent Sandboxes to the Cloud, Billed by the Second From $0.07 an Hour

Docker's new Cloud Sandboxes put its microVM agent isolation on hosted compute, billed by the second from $0.07 an hour, and hand the Kits format to the CNCF.

Seung Jung7 days ago
A Million Agent Skills in Seven Months — and Nearly Half Were Installed Exactly Once
Developer Tools

A Million Agent Skills in Seven Months — and Nearly Half Were Installed Exactly Once

Vercel's skills.sh registry hit 1M agent skills in seven months. Nearly half were installed once, while 375 listings account for 62% of all installs.

Seung Jung6 days ago
AWS Open-Sources Strands Harness — and Names the One Rival That Undercut It
Developer Tools

AWS Open-Sources Strands Harness — and Names the One Rival That Undercut It

The Strands Agents team at AWS published Strands Harness on September 21 under Apache 2.0, for Python and TypeScript, deployable on a laptop or on any of five c...

Seung Jung10 days ago
Zed's Delta Replaces Pull Requests With Agent Threads
Developer Tools

Zed's Delta Replaces Pull Requests With Agent Threads

Zed's Delta swaps the pull request for a shared thread where agents and reviewers work from the same context. Its own repo already has PRs off.

Seung Jung11 days ago
832,378 Lines of Rust in 14.5 Weeks: Inside an Agent-Run Rewrite
Developer Tools

832,378 Lines of Rust in 14.5 Weeks: Inside an Agent-Run Rewrite

GitHub converted 430,000 lines of TypeScript into 832,378 lines of Rust in 14.5 weeks using coding agents. Memory use fell from 1,383MB to 126MB.

Seung Jung13 days ago