Decision model guide

Cloudflare Clef. Context in, decisions out.

Clef evaluates a shared state against questions you define, returning probabilities over allowed answers. Use it for routing, classification and checks inside an application.

A guide to Cloudflare's model. The example below runs on Workers AI with your Cloudflare account.

One state, three decisions

Incoming support message

“My account is locked after too many sign-in attempts. I need access before my meeting.”

choice

Choose the responsible team

score

Assess urgency on an ordered scale

noul

Estimate whether access is blocked

Illustrative questions; no live inference is performed here.

Clef parameters
27B
Workers AI context tokens
65,536
Per million input tokens
$0.24
Open weights license
Apache-2.0

Workers AI specifications and pricing checked on October 3, 2026. See the official documentation for updates.

01 / How it works

Define the possible answers before you ask

Clef combines a Qwen backbone with a joint schema head. One forward pass scores the valid options across questions, without generating an intermediate text response.

choice

Pick a category

Supply named options with descriptions. Read the selected option, confidence and per-option probabilities.

score

Evaluate on a scale

Supply ordered criteria. The answer includes a probability-weighted score over zero-based levels.

noul

Ask a yes/no question

Read the probability of true. Set your application's threshold and decide when a person should review the result.

02 / Clef and Clef-flash

Two sizes for different latency budgets

Cloudflare positions Clef for precision and Clef-flash for time-sensitive decisions. Evaluate both against representative examples from your own workflow.

ModelBackboneReported median latency
ClefQwen3.8-27B209.3 ms
Clef-flashQwen3.5-9B38.8 ms

Latency figures are from Cloudflare's published internal Decision Index evaluation. They are not a production latency guarantee.

Read the launch and evaluation details

03 / Workers AI quick start

Send state. Define questions. Read answers.

Add an AI binding named AI to your Worker, then call @cf/cloudflare/clef. The example evaluates three aspects of the same support message.

Request shape

Pass model, state and questions. Workers AI accepts 1–64 questions and returns answers under the same question IDs.

Images on Workers AI

The hosted API accepts up to four embedded PNG, JPEG or WebP images; remote image URLs are not accepted. Check the reference for payload limits.

Jev / SystemOne compatibility

Clef follows the typed request and answer format. When switching providers, update authentication and the endpoint, and check the provider's response envelope.

Open the Workers AI reference
Worker / TypeScript
export default {
  async fetch(request, env) {
    const result = await env.AI.run("@cf/cloudflare/clef", {
      model: "clef",
      state: "My account is locked after too many sign-in attempts.",
      questions: {
        team: {
          type: "choice",
          instructions: "Which team should help?",
          criteria: {
            accounts: "Sign-in and account access",
            billing: "Payments and invoices",
            technical: "Service errors and outages"
          }
        },
        urgency: {
          type: "score",
          instructions: "How urgently does this need attention?",
          criteria: ["Low", "Medium", "High"]
        },
        blocked: {
          type: "noul",
          instructions: "Is account access blocked?"
        }
      }
    });

    return Response.json(result);
  }
} satisfies ExportedHandler<{ AI: Ai }>;

04 / Run it your way

Hosted inference or your own deployment

Open weights

The Hugging Face release includes the backbone and joint schema head. Follow its load_release_model and systemone examples for self-hosting, including local image and video inputs.

Read the model card

Common questions

Before you integrate Clef

Does Clef generate chat replies?

Clef returns decisions constrained to your schema. Use those answers to choose the next action; use a text-generating model when your application needs a written reply.

Are the hosted and local input limits identical?

No. Workers AI lists a 65,536-token context window and documents embedded image inputs. The local model card also describes video inputs, and encode_record defaults to 16,384 tokens. Follow the documentation for the deployment you use.

Can I use Clef commercially?

The weights are published under Apache-2.0. Review the included license and notices for self-hosting, and Cloudflare's service terms for hosted use.

How should I use the returned probabilities?

Choose thresholds using labelled examples from your workflow. Measure incorrect actions and missed cases, and route uncertain decisions to a fallback or a reviewer.