Robots Atlas>ROBOTS ATLAS
Artificial Intelligence

Cloudflare's Clef: a 38.8 ms decision and Jev API compatibility

Sir Robot11 October 2026 · 2 min read
Cloudflare's Clef: a 38.8 ms decision and Jev API compatibility

Cloudflare has released Clef and Clef-flash, decision models that return probabilities for schema options instead of generating text. Both landed on 1 October 2026: weights on Hugging Face under Apache 2.0, hosting on Workers AI, and an API compatible with Jev. The flash variant has a median latency of 38.8 ms.

Key takeaways

  • Clef is 27B on a Qwen3.8-27B backbone with a vision encoder, Clef-flash is 9B on Qwen3.5-9B
  • A 64k-token context window, accepting text, JSON, images and video
  • Median latency: 38.8 ms for flash, 209.3 ms for Clef, 524.1 ms for Jev
  • Top of the Jev Decision Index and ahead across 43 benchmarks, per Cloudflare
  • Apache 2.0 licence, weights on Hugging Face, hosting on Cloudflare Workers AI
38.8 msmedian latency for Clef-flash — 13 times faster than JevCloudflare, 1 Oct 2026

One prefill pass, no generation

Clef is not Autoregressive model: One that builds its answer token by token, each time appending to what it has already said. Every token is a separate forward pass.. The model runs a single prefill pass over the state and the schema, then scores every permitted field value in parallel. Dozens of decoding steps collapse into one forward pass — hence the latency gap.

Input
State + question schema
Model
One prefill pass
Stage 1: options pull context from the state
Stage 2: fields cross-attend
Parallel scoring of every optionAllow
Schema-aligned probabilities

The second piece is two-stage attention routing: the permitted options first pull the relevant context out of the state, then the fields can see each other before scoring. In practice the model decides several fields at once while keeping them consistent.

Two variants, two trade-offs

SpecClefClef-flash
Size27B9B
BackboneQwen3.8-27BQwen3.5-9B
Median latency209.3 ms38.8 ms
Multimodal inputtext, JSON, image, videotext, JSON, image, video
LicenceApache 2.0Apache 2.0

API compatibility as leverage

The biggest advantage is not the architecture but the interface. Clef is fully compatible with the Jev API, so a team on TypeSafe AI's closed model swaps the endpoint and keeps its own code. It is the same playbook clouds used to pull traffic away from the OpenAI API.

Four parameters that decide a deployment:

Qwen3.8-27Bbase model for the Clef variant
64k ctxcontext window
bf16 safetensorsweight format
Jev-compatible APIendpoint swap, no code changes

The base is Qwen3.8-27B, with weights on the Hugging Face Hub — the third decision model this week built on open Qwen weights.

Cloudflare adds a pledge about requests: we don't read, store, or train on them. That is a contractual promise rather than a technical property of the model — read it alongside the Workers AI terms.

Why it matters

The decision layer in agents is high-frequency, low-value-per-call traffic — exactly what a CDN monetises. Cloudflare is not selling a model here so much as the network edge. Open weights are an argument against vendor lock-in rather than a product in themselves.

What's next

  • A self-serve RL platform for fine-tuning is announced but has no release date
  • Custom fine-tuning currently runs through Cloudflare's forward-deployed engineering team
  • The 43-benchmark results come from the vendor, with no independent verification

Sources

Share this article