Cloudflare has released Clef and Clef-flash, decision models that return probabilities for schema options instead of generating text. Both landed on 1 October 2026: weights on Hugging Face under Apache 2.0, hosting on Workers AI, and an API compatible with Jev. The flash variant has a median latency of 38.8 ms.
Key takeaways
- Clef is 27B on a Qwen3.8-27B backbone with a vision encoder, Clef-flash is 9B on Qwen3.5-9B
- A 64k-token context window, accepting text, JSON, images and video
- Median latency: 38.8 ms for flash, 209.3 ms for Clef, 524.1 ms for Jev
- Top of the Jev Decision Index and ahead across 43 benchmarks, per Cloudflare
- Apache 2.0 licence, weights on Hugging Face, hosting on Cloudflare Workers AI
One prefill pass, no generation
Clef is not autoregressive?Autoregressive model: One that builds its answer token by token, each time appending to what it has already said. Every token is a separate forward pass.. The model runs a single prefill pass over the state and the schema, then scores every permitted field value in parallel. Dozens of decoding steps collapse into one forward pass — hence the latency gap.
The second piece is two-stage attention routing: the permitted options first pull the relevant context out of the state, then the fields can see each other before scoring. In practice the model decides several fields at once while keeping them consistent.
Two variants, two trade-offs
| Spec | Clef | Clef-flash |
|---|---|---|
| Size | 27B | 9B |
| Backbone | Qwen3.8-27B | Qwen3.5-9B |
| Median latency | 209.3 ms | 38.8 ms |
| Multimodal input | text, JSON, image, video | text, JSON, image, video |
| Licence | Apache 2.0 | Apache 2.0 |
API compatibility as leverage
The biggest advantage is not the architecture but the interface. Clef is fully compatible with the Jev API, so a team on TypeSafe AI's closed model swaps the endpoint and keeps its own code. It is the same playbook clouds used to pull traffic away from the OpenAI API.
Four parameters that decide a deployment:
The base is Qwen3.8-27B, with weights on the Hugging Face Hub — the third decision model this week built on open Qwen weights.
Why it matters
The decision layer in agents is high-frequency, low-value-per-call traffic — exactly what a CDN monetises. Cloudflare is not selling a model here so much as the network edge. Open weights are an argument against vendor lock-in rather than a product in themselves.
What's next
- A self-serve RL platform for fine-tuning is announced but has no release date
- Custom fine-tuning currently runs through Cloudflare's forward-deployed engineering team
- The 43-benchmark results come from the vendor, with no independent verification
Sources
- Cloudflare Blog — Clef decision models
- Hugging Face — Cloudflare/clef
- Hugging Face — Cloudflare/clef-flash





