Cloudflare Releases Two AI Models That Pick Answers Instead of Writing Them

Cloudflare released two decision models on Oct. 1, Clef and the smaller Clef-flash, and published the weights for both on Hugging Face under the Apache 2.0 license. Both also run as a hosted service on the company's own network, through its Workers AI platform.
A decision model answers a fixed set of questions about an input instead of writing free text. Cloudflare's example is a support message: the model returns probabilities for how urgent it is and which team should handle it, and the surrounding code can then route the ticket, escalate it or pass it to a person. Cloudflare says its threat-intelligence team has been using Clef to sort website domains into categories.
Both models start from existing Qwen models, Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash. The company left those unchanged and trained a smaller layer on top to do the choosing. Cloudflare says Clef makes a single pass over the input and then scores all the allowed answers at once rather than writing one out word by word, and credits that design for the speed.
The speed and accuracy comparisons in Cloudflare's announcement are the company's own figures, measured against models it chose, and no independent evaluation has been published. Its table gives Clef a median response time of 209.3 milliseconds and Clef-flash 38.8, against 524.1 for Jev, the Typesafe AI decision model that Clef is designed to replace without a code change. The accuracy results are mixed: on two of the ten benchmarks in the same table, Jev scores higher than both Clef models.
Cloudflare also opened a service to fine-tune Clef on a customer's own data using reinforcement learning, run at first by its forward-deployed engineering team and offered later as a self-serve product. The company says it does not read, store or train on requests to the hosted models unless the customer is using that fine-tuning service.
