Cloudflare released Clef and Clef-flash on 1 October 2026: two open-weight decision models that return a probability for every allowed answer instead of generating text. Both are Apache 2.0, both read images (Jev, the model they copy the API of, reads only text), and both are hosted on Workers AI at $0.24 and $0.09 per million input tokens. Cloudflare says Clef "is currently the leader" on the community Decision Index. We rebuilt that index's scoring formula, ran Cloudflare's own numbers through it, and read the weight files byte by byte. The leader claim holds on paper. Four details in the release matter more.

The short version: Clef would top the board by 3.3 points, but the score is self-reported and two benchmarks are missing from the table. Clef costs 5.7 times Jev per token. Clef-flash falls below random guessing on hallucination detection. And as of 2 October, 5,354 downloads went to GGUF repackages that contain no decision head at all.

What Cloudflare shipped

A decision model takes a "state" (text, JSON or images) plus up to 64 typed questions and returns calibrated probabilities: noul for yes or no, choice for named options, score for ordered levels. There is no free-form output to parse. Clef sits on Qwen3.8-27B, Clef-flash on Qwen3.5-9B, and both add a small "joint schema head" that scores every option of every question in one forward pass. The Clef model card lists the files and says it was tested "on a single H200".

ClefClef-flashJev (TypeSafe)
Base modelQwen3.8-27BQwen3.5-9BClosed
Parameters (backbone)27.36 billion9.41 billionNot disclosed
Download size (bf16)55.0 GB19.1 GBNo weights
Hosted price, per 1M input tokens$0.24$0.09$0.042
Context window (Workers AI)65,536 tokens65,536 tokens32K (per Cloudflare)
Image inputUp to 4 per requestUp to 4 per requestNo
LicenseApache 2.0Apache 2.0API only

The hosted limits come from the Workers AI model page: images must be embedded as base64 (PNG, JPEG or WebP, 4 MiB and 16 megapixels each, 8 MiB total), and remote image URLs are rejected. Jev's price is from TypeSafe's homepage, which quotes "$42 per billion input tokens".

We rescored the leader claim

The Decision Index is a community leaderboard for Jev-style models, maintained in a public Hugging Face Space. Its 0.2.1 edition averages 38 benchmarks in five areas, chance-corrects each score, and weights arts at 10% and the other four areas by the square root of their benchmark count. Clef is not on it: the Space's last update (2 October, 01:22 UTC) renamed another entrant, and its data files do not mention Clef. The numbers in Cloudflare's model card come from what it calls "our internal run" of the suite.

So we reimplemented the published formula and checked it against all 71 entrants on the board. It reproduces every published score to two decimals. Then we fed in Cloudflare's numbers for Clef and Clef-flash.

ModelIndex 0.2.1 scorePlace on the 71-entrant board
Clef (Cloudflare's numbers, 2 missing benchmarks scored as zero)61.21Would be 1st
Clef (missing benchmarks at the best score on the board)62.72Would be 1st
Jev (independent run)57.911st today
Surogate Rune 26B-A4B v357.442nd today
Decider chat, Gemma-4-31B57.333rd today
Clef-flash (missing benchmarks scored as zero)57.07Would be 4th

Two of the 38 index benchmarks are absent from Cloudflare's table: HLE and iSarcasmEval. Cloudflare's blog says it ran "43 eval benchmarks" and its model card publishes 41 rows. Neither gap changes Clef's rank. HLE barely moves anyone (the best skill score on the board is 4.7 out of 100, Jev's), so even scoring both as zero leaves Clef 3.3 points ahead. Clef-flash is different: at zero it would place fourth, and with Jev's scores on the two missing benchmarks it would edge into first, so its position depends on numbers nobody has published.

Cloudflare did not adjust its competitors. Every Jev, Kev 9B, Laya and DiffusionGemma figure in its 41-row table matches the independent board exactly, latency included. What remains unverified is Clef's own column. The index counts rows an entrant trained on as wrong, and Cloudflare trained on "internal synthetic datasets" that nobody outside the company can check for overlap.

Decision Index 0.2.1 score from Cloudflare's own numbers: Clef 61.2 against Jev's independently measured 57.9
Clef would lead the Decision Index on Cloudflare's own numbers, 61.2 to Jev's 57.9. Not yet verified independently.

Where Clef loses to Jev

A higher average hides a pattern. On the 36 index benchmarks Cloudflare published, Clef beats Jev on 25 and loses on 11, and the losses cluster in multi-step reasoning, where Clef trails by double digits:

BenchmarkClefClef-flashJevClef vs Jev
GPQA Diamond (graduate science)48.051.078.3-30.3
BBH (hard reasoning)73.768.992.9-19.2
MMLU-Pro65.965.382.7-16.8
When2Call (should the agent call a tool?)72.465.681.0-8.6
Home appliance simulator83.097.752.3+30.7
Habermas Machine68.771.845.9+22.8
CLadder (causal questions)94.097.772.6+21.4
BANKING77 (intent routing)94.290.979.7+14.5

The practical reading: Clef is stronger at routing and labelling (intents, products, tool selection, structured checks) and weaker when the right answer needs knowledge or several reasoning steps. When2Call matters most for agent builders: it measures whether a model knows to call a tool, ask a question or decline, and Jev is ahead there by 8.6 points. Cloudflare's latency row (Clef 209 ms median, Jev 524 ms) needs a footnote: the index labels Jev's figure a hosted-API network round trip, "not comparable to the on-card single-process figures", and Cloudflare does not say how it timed Clef.

Clef-flash is not a smaller Clef

Cloudflare positions Clef-flash as the option for latency-critical decisions, at 38.8 ms median. On most rows it tracks Clef closely or beats it. On two it collapses, and both are checks people put in production:

  • RAGTruth (spotting hallucinated answers): Clef scores 79.4, Jev 76.5, Clef-flash 35.6. The index's random baseline for that benchmark is 51.8, so Clef-flash scores below chance and earns zero skill.
  • CLINC150 with out-of-scope requests: Clef 97.4, Jev 89.3, Clef-flash 66.8. This is the test of saying "none of the above" when a request fits no category, a 30.6-point drop from the larger model.

If your pipeline uses a decision model as a guardrail (is this answer grounded, does this request belong to any of my categories), Clef-flash's speed advantage costs you the guardrail. Use Clef, or Jev, for those two jobs.

RAGTruth hallucination detection: Clef scores 79.4, Clef-flash 35.6, below the 51.8 random baseline
RAGTruth hallucination detection: Clef 79.4, Clef-flash 35.6, below the 51.8 chance line.

What is in the 55 GB download

Cloudflare's blog says it froze the Qwen backbone and trained the head "alongside rank-256 low-rank adapters". We checked what that means in the published files by reading tensor headers and byte ranges straight from Hugging Face (HTTP range requests, no full download) and comparing them with the base models, Qwen3.8-27B and Qwen3.5-9B.

  • Clef: layers 0 to 39 of its 64 language layers are byte-identical to stock Qwen3.8-27B in every sample we took. Layers 40 to 63 differ in every sample. The vision encoder blocks and the output head we sampled are unchanged. The adapters appear to have been merged into the top 24 layers only.
  • Clef-flash: every language layer we sampled (0, 3, 15 and 31) differs from Qwen3.5-9B. Its vision blocks are unchanged.
  • The joint head is a separate 256 MB file, joint_head.safetensors, holding 128 million parameters in four small transformer layers. It is what turns hidden states into option scores.

The image-reading part of Clef is therefore stock Qwen vision feeding adapted top layers. That matters for creative use: Clef classifies images as well as Qwen sees them, and Cloudflare published no image-specific benchmark. Test it on your own assets before trusting it.

Most local downloads are missing the decision head

Within a day, Hugging Face listed more than 25 community repackages. We checked each for joint_head.safetensors. None of the ten GGUF repos we found, the format llama.cpp and most desktop apps load, includes it, and neither do two FP8 conversions. The five GGUF repos with downloads:

RepositoryFormatDecision head includedDownloads (2 Oct, 16:12 UTC)
bartowski/Cloudflare_clef-flash-GGUFGGUFNo3,011
abenzerps/Clef-GGUFGGUFNo810
bartowski/Cloudflare_clef-GGUFGGUFNo772
prithivMLmods/clef-flash-GGUFGGUFNo487
prithivMLmods/clef-GGUFGGUFNo274
Cloudflare/clef-flash (official)bf16Yes1,303
Cloudflare/clef (official)bf16Yes824
mlx-community/clef-flash-4bitMLX 4-bit, 6.2 GBYes215

That is 5,354 downloads of files without the head against 2,127 for Cloudflare's own repos. A GGUF of Clef loads as a Qwen-architecture text model with modified top layers, and the scoring code (joint_schema_model.py) and head weights are not in the file. The mlx-community model card spells out what happens: "LM Studio will load the backbone but produce meaningless text." Packages that do carry the head include the mlx-community 4-bit builds (6.2 GB for Clef-flash, 16.3 GB for Clef, run with their bundled clef_mlx.py loader), simonlehmann's NVFP4 Clef (24.1 GB) and ramgpt's EXL3 Clef-flash (8.4 GB). Ollaya's Clef package pulls Cloudflare's own Clef-flash weights by hash and reports identical decisions to Cloudflare's code on 571 questions from 131 requests.

What it costs per 1,000 decisions

All three hosted models list a price for input tokens only, so cost scales with how much state you send. Cloudflare's page does not say how images are counted; each response carries a usage field, so measure one image request before you budget a batch.

Input per decisionJev ($0.042/M)Clef-flash ($0.09/M)Clef ($0.24/M)
500 tokens (a comment or short message)$0.021$0.045$0.12
2,000 tokens (a long brief or ticket)$0.084$0.18$0.48
10,000 tokens (a full document)$0.42$0.90$2.40

Clef is 5.7 times Jev per token and Clef-flash 2.1 times. At creator volumes (10,000 comment triages a month at 500 tokens each) that is $0.21 on Jev against $1.20 on Clef, so the price only matters at scale. What Clef buys for the difference is image input, an open license you can self-host, and a 64K context.

Hosted price per million input tokens: Jev $0.042, Clef-flash $0.09, Clef $0.24
Price per million input tokens: Jev $0.042, Clef-flash $0.09, Clef $0.24.

How to use Clef for creative work

Clef's one capability Jev lacks is reading images, so the creator use cases are visual triage: sorting a folder of generated images, checking thumbnails against a brief, flagging user submissions before a human looks. A workable setup:

  1. Pick the job by the tables above. Labelling, routing and yes-or-no checks: Clef or Clef-flash. Hallucination checks or "none of the above" detection: Clef, not Flash. Questions that need real reasoning: neither, use an LLM.
  2. Write each check as a typed question. One noul per rule ("Does the image contain readable text?"), one choice per category set, one score per graded judgment. Up to 64 questions share one request, so a full brief costs one call.
  3. Embed the image. Send it as a base64 data URL (data:image/png;base64,...) in the images array, at most four per request and 16 megapixels each. Downscale first: you pay for every input token.
  4. Call the hosted model. The endpoint is /ai/run/@cf/cloudflare/clef on your Cloudflare account, with "model": "clef" or "clef-flash" in the body. Because the API matches Jev's, code written for Jev needs only the URL and model name changed.
  5. Act on the probability, not the label. Auto-approve above a threshold, send everything else to a human. Calibrate on 50 to 100 of your own labelled examples before you set it.
  6. Self-host from a package that has the head. Use Cloudflare's repos with joint_schema_model.py, an MLX build that lists joint_head.safetensors, or Ollaya. Skip the GGUF files.
  7. Budget the hardware. Clef in bf16 is 55 GB and needs an 80 GB-class GPU; Clef-flash at 19.1 GB should fit a 24 GB card. On a Mac, the 6.2 GB MLX Clef-flash leaves room on a 16 GB machine.

A sample request body for a thumbnail check, following the documented schema (we have no Workers AI account on this machine, so this was not executed):

{
  "model": "clef",
  "images": ["data:image/jpeg;base64,..."],
  "state": "Candidate thumbnail for a video titled 'Three lighting setups for product shots'.",
  "questions": {
    "has_text": { "type": "noul", "instructions": "Does the image contain readable text?" },
    "subject": {
      "type": "choice",
      "instructions": "What is the main subject?",
      "criteria": { "product": "A product on a surface", "person": "A person", "scene": "A room or landscape" }
    },
    "clutter": { "type": "score", "criteria": ["Clean", "Some clutter", "Busy"] }
  }
}

Frequently asked questions

What is Cloudflare Clef?

Clef is an open-weight decision model released by Cloudflare on 1 October 2026. Instead of generating text, it returns a probability for each allowed answer to typed questions about text, JSON or images. It comes in two sizes, Clef (27B) and Clef-flash (9B), under Apache 2.0.

Is Clef better than Jev?

On Cloudflare's own numbers, Clef would lead the Decision Index at 61.2 against Jev's 57.9 under the index's published formula. It beats Jev on 25 of 36 published index benchmarks but trails by 17 to 30 points on GPQA Diamond, BBH and MMLU-Pro. The Clef scores have not been reproduced independently.

How much does Clef cost?

On Workers AI, Clef costs $0.24 and Clef-flash $0.09 per million input tokens, against $0.042 for Jev. The weights are free to download and self-host under Apache 2.0.

Can I run Clef in LM Studio or llama.cpp?

Not as a decision model. None of the GGUF repackages we checked includes the joint schema head that produces Clef's probabilities; LM Studio loads the backbone and produces meaningless text. Use Cloudflare's own weights with its Python code, an MLX build that includes joint_head.safetensors, or Ollaya.

Can Clef classify images?

Yes. The hosted API accepts up to four embedded PNG, JPEG or WebP images per request, and the open weights also accept video frames. Its vision encoder is unchanged from Qwen, and Cloudflare published no image benchmark, so test it on your own images first.

Should I use Clef or Clef-flash?

Clef-flash is about five times faster at the median and cheaper, and it matches or beats Clef on many routing tasks. Avoid it for hallucination checks (35.6 on RAGTruth, below chance) and for detecting requests that fit no category (66.8 against Clef's 97.4 on CLINC150).