Back to all articles

API & Integration

What Is Decisions API? OpenAI's Fast Decision Layer Explained

What is Decisions API? Learn how OpenAI's limited-preview decision layer handles bounded classification, routing, and agent next steps—and how to evaluate it safely.

By Jev AISep 30, 202612 min read
What Is Decisions API? OpenAI's Fast Decision Layer Explained

What Is Decisions API? OpenAI's Fast Decision Layer Explained

If you are searching for what is Decisions API, the short answer is this: OpenAI Decisions API is a limited-preview interface for making small, bounded judgments inside software. Instead of asking a general-purpose model to write a paragraph and then parsing that paragraph, a developer defines a question, supplies context, and lets the system choose from a finite set of answers.

That makes the API interesting for classification, request routing, tool-call gates, and the next step in an AI agent loop. It is not a replacement for a general chat or reasoning model. It is a narrower layer that can sit between application state and deterministic code.

OpenAI presented the Decisions API at DevDay 2026. Contemporary launch coverage describes a specialized version of GPT-6 Luna, text or image context, predefined answers, and a limited-preview launch. The reports also mention an answer time of about 150 milliseconds and a roughly tenfold speed advantage over a regular GPT-6 Luna call. Those figures are reported launch claims, not a service-level guarantee.

This guide explains the idea without pretending that third-party launch coverage is a stable API reference. The preview contract, pricing, limits, and access rules may change. Use the architecture and evaluation principles here as a practical starting point, then verify the current OpenAI documentation and dashboard before shipping a critical workflow.

Table of contents

What is Decisions API in one sentence?

Decisions API is a decision-oriented model interface that takes application context and a bounded question, then returns a structured choice that software can use.

The word bounded is the key. “Write a helpful answer to this customer” is an open-ended generation task. “Which approved queue should handle this ticket?” has a finite answer space. “Does this exact tool call need human confirmation?” is also bounded if the application defines what “needs confirmation” means.

A useful abstraction is:

decision = f(state, question, allowed_answers)

The output is not meant to be a polished essay. It is a signal for the surrounding program. The program still validates the response, checks permissions, applies business rules, records the outcome, and decides whether an action is allowed.

Public launch reporting says the preview targets fast classification, routing, and agent next-step decisions, with text or image context and a limited list of possible answers. Treat the request and response schema as provisional until OpenAI updates its official reference.

How is a decision API different from a chat API?

Most model integrations are chat-shaped: messages go in and free-form text comes out. That flexibility is valuable when the product needs explanation, drafting, synthesis, or open-ended reasoning. It is less convenient when the model sits in a tight control loop that runs thousands of times a day.

Imagine a support system receiving this ticket:

The customer was charged twice and payouts have failed for three days.

A general model may produce a useful paragraph that mentions billing, payments, and urgency. The application must then extract a route, validate the label, deal with extra wording, and decide what to do if the format drifts. A bounded question starts with the software contract instead:

Question: Which approved queue owns this ticket?
Answers: billing, technical, account, none_of_the_above
Context: <ticket state>

The difference is not merely “JSON versus text.” A structured-output feature can ask a generative model to format several fields as JSON. A decision API is designed around the semantic choice itself: the answer space is declared before inference, and the result is intended to be consumed as a decision signal.

The two patterns can be simplified like this:

chat model:       prompt -> prose -> parser -> validation -> retry -> action
decision model:   state + bounded question -> decision -> policy -> action

The second path is smaller, but it is not risk-free. A valid answer can still be wrong. A confidence number can still be miscalibrated. A model should not be allowed to grant a permission just because it selected “safe.” The advantage is a clearer boundary: model judgment on one side, deterministic authority on the other.

A mathematical sketch of loose model output passing through a validation sieve into a typed decision

The practical design rule

Use a decision API when the application can write down the answer space before the call. Use a general model when the application needs the model to generate language, explore possibilities, or make a plan that cannot be reduced to a small set of approved outcomes.

How does the decision loop work?

A reliable integration can be designed as four steps.

1. Capture only the state needed for one decision

State is the evidence available to the decision. It might be a ticket, an email, a proposed tool call, a screenshot, a compact object from several services, or a short list of retrieved facts.

Keep it focused. Sending an entire conversation archive when a route only needs a ticket, account tier, and recent events adds noise, latency, privacy exposure, and cost. Ask what a human reviewer would need to answer this one question.

2. Write one bounded question

The question should describe one judgment. “Classify, prioritize, refund, and notify the customer” hides several policies. Split it into separate questions or application steps:

  • Which approved team should own the case?
  • What is the operational severity on a defined scale?
  • Does the next action require human approval?

Each question should have an answer space that a reviewer can inspect. Include an explicit fallback such as none_of_the_above when the real world may not fit the listed categories.

3. Define the allowed answers

The application should own the options: billing, technical, account, and manual_review for routing; allow, confirm, and block for a tool gate. For priority, use a rubric with operational definitions rather than vague labels such as “low” and “high.”

The question design is part of the product policy. Changing the options changes the meaning of historical results, so version the question, rubric, and policy alongside the model identifier.

4. Apply policy before code acts

The returned decision is a signal. Your application must still check permissions, resource scope, data validity, and side-effect rules. A model can recommend refund_review; only an authorized service or human should approve an actual refund.

A hand-drawn mathematical sketch of state, a bounded question, finite answers, and a structured result

The safest equation is:

executable path = model judgment ∩ deterministic policy

The intersection is the point. A decision model handles an ambiguous semantic judgment; code keeps ownership of authority.

What can you use Decisions API for?

The best use cases are repeated, narrow decisions where the next action already exists in the surrounding system.

Classification and routing

Route tickets, leads, documents, incidents, or moderation events to an approved queue. The route can be based on intent, product area, severity, customer tier, or a combination of facts in the state. A fallback route prevents the system from forcing every unfamiliar case into a misleading category.

Choosing an agent's next step

An agent can use a stronger model to understand the goal and plan the task, then use a fast decision layer to choose the next step from an allowlist: search, open, fill, verify, retry, ask for help, or finish. The host application still checks whether the selected action is available and whether its arguments are safe.

Tool-call and transaction gates

Before sending an email, changing an account, deleting data, making a payment, or publishing content, ask a narrow question about the proposed action. Combine the answer with deterministic allowlists, user permissions, confirmation requirements, and audit logging. The model should help interpret intent; it should not be the only authorization layer.

Model routing and cost control

Not every request needs a frontier model. A decision layer can route simple work to a fast classifier, difficult work to a stronger model, uncertain work to retrieval, or sensitive work to a person. The route should be selected from capabilities with known input shapes, latency bands, costs, and fallbacks—not from arbitrary model names hidden in a prompt.

A mathematical sketch of a decision boundary routing requests to fast, deep, retrieval, or human paths

Triage and prioritization

Queueing systems often need a consistent priority signal, not a written summary. A score is useful when the rubric says what each level means operationally. “Service blocked” and “minor inconvenience” are more testable than “urgent” and “not urgent” without definitions.

Visual decisions

The launch coverage describes image context as part of the preview. That could support decisions such as identifying a known UI state from a screenshot or routing an image-based report. Treat this as a capability to verify, not a reason to assume every image workflow is ready for production. Check supported formats, size limits, privacy handling, and accuracy on your own labeled examples.

For a public decision workflow, a useful implementation pattern is to expose state, typed questions, and structured results in an interactive form. That pattern is useful for understanding the category, but it is not an OpenAI-compatible endpoint.

What might an integration look like?

Because the preview contract may change, do not copy an unofficial payload into production. Put the provider behind a small adapter and treat the following as conceptual pseudocode:

{
  "context": {
    "ticket": "The customer was charged twice.",
    "account_tier": "business",
    "recent_events": ["payment_succeeded", "payment_succeeded"]
  },
  "question": {
    "name": "route",
    "prompt": "Which approved workflow owns this case?",
    "answers": ["refund_review", "technical_support", "account_security"]
  }
}

The adapter should validate the response as unknown, reject options outside the allowlist, and separate uncertainty from transport failure. A timeout is not a confident “no.”

const decision = await decisionsApi.evaluate(request);

if (!allowedRoutes.includes(decision.choice)) {
  return sendToManualReview('Unknown route');
}

if (decision.score < ROUTE_THRESHOLD || !policyAllows(decision.choice)) {
  return sendToManualReview('Uncertain or disallowed route');
}

return dispatch(decision.choice, { auditId, source: 'decisions-api' });

The field names are illustrative. Do not assume score or confidence is calibrated. Record the raw result, question version, policy version, model identifier, final action, and later human outcome so you can measure the threshold.

Other decision systems use a similar state-plus-typed-questions pattern, including Choice, Score, and yes/no-style decisions. Those are separate product contracts, not an OpenAI Decisions API specification.

OpenAI Decisions API versus Jev

OpenAI's preview and Jev address a similar category: fast, structured decisions inside software. They are not the same service, and the available information describes different contracts.

Dimension OpenAI Decisions API Jev AI
Status described by current coverage Limited preview at DevDay 2026 Public playground and API workflow
Reported engine Specialized version of GPT-6 Luna Jev decision model
Input described publicly Text or image context Text, JSON objects, and arrays of text on the public site
Output idea Selection from developer-defined answers, with a score reported in coverage Typed Choice, Score, and Noul results with probabilities and confidence
Example jobs Classification, routing, and agent next step Classification, routing, scoring, safety checks, and review gates
Contract maturity Verify endpoint, limits, pricing, and access in the preview Public docs and interactive playground available

The useful comparison is “which contract can my team evaluate and operate?” OpenAI may be attractive when account consolidation, image context, or preview access matters. Jev may be attractive when a team wants a public playground and documented typed decisions. In either case, define a small answer space, collect labeled examples, and keep authority in application code.

The complementary architecture is straightforward: give a system state, ask typed questions, and let your own logic route, queue, block, or request review. That pattern can inform an OpenAI adapter without implying feature parity.

Where does it fit in an AI agent?

A production agent usually has several distinct layers:

  1. Orchestrator: tracks the task, context, retries, and loop state.
  2. Generative model: interprets intent, plans work, writes language, or summarizes evidence.
  3. Decision layer: answers narrow questions about route, risk, priority, or completion.
  4. Policy and permissions: enforce what the user, agent, and tool are allowed to do.
  5. Action and audit layer: executes approved calls and records what happened.

The Decisions API belongs in layer three. It should not silently become layer four. A decision model may say that a tool call appears low-risk; it cannot grant a permission that the policy engine did not grant.

A mathematical sketch of sequential safety gates before an AI agent can execute an action

For a first experiment, select one reversible, low-risk decision:

historical examples
        ↓
question + allowed answers
        ↓
offline evaluation
        ↓
shadow traffic
        ↓
human-approved automation
        ↓
monitored production path

Use shadow mode for money movement, access control, deletion, safety, and reputation. Compare the decision with a human label or trusted rule before it can change the world.

Latency, cost, and confidence

Latency

The launch reports cite roughly 150 milliseconds and about ten times the speed of a regular GPT-6 Luna call. Measure p50, p95, and p99 yourself; network distance, context size, concurrency, retries, and queueing can dominate model time.

Cost

A decision endpoint may reduce cost by avoiding long explanations and routing easy work away from expensive models. Include context, retries, observability, downstream calls, and human review in unit economics. Verify preview pricing in the current OpenAI account.

Confidence and calibration

A score is evidence, not permission. A result of 0.92 can still be wrong on a specific language, segment, or adversarial input. Measure:

  • agreement with reviewed labels;
  • false positives and false negatives for each costly action;
  • automation coverage at each threshold;
  • calibration by class, language, input type, and customer segment;
  • human-review volume and time to resolution;
  • drift after a model, question, policy, or product change.

If the best and second-best options are close, the top label may be fragile even when it looks decisive. Keep an abstain or review path. For high-impact actions, use a two-key design: the decision model recommends a path, and deterministic policy or a human authorizes it.

A side-by-side mathematical sketch contrasting open-ended generation with bounded typed decisions

What should developers verify before production?

Public sources describe Decisions API as a limited preview, so verify these details before integrating deeply:

  1. the official endpoint, authentication scope, and current request schema;
  2. whether image input is enabled for your account and supported use case;
  3. how the answer set is defined and whether multiple questions are supported;
  4. what the returned score means and whether it is calibrated or only model-reported;
  5. limits, latency expectations, rate-limit behavior, retries, and error codes;
  6. data retention, privacy controls, and regional availability;
  7. the billing unit, including failed, repeated, or batched calls;
  8. model-version pinning and the process for preview changes.

Keep the integration behind a narrow provider adapter. Version questions and policies, redact sensitive state from logs, and treat user content as data—not policy. If the provider is unavailable, route to a safe queue or human review rather than guessing.

Frequently asked questions

Is Decisions API a replacement for ChatGPT or a general Responses API?

No. It is better understood as a specialized decision layer. Use a general model for generation, planning, tool orchestration, explanation, and user-facing text. Use a decision endpoint when the application needs a narrow judgment from an approved set of outcomes.

Does Decisions API return text?

Public coverage emphasizes selecting from developer-defined answers rather than generating a free-form paragraph. The transport may be structured, but the useful output is the selected decision and associated metadata—not an essay for a user to read.

Can it choose the next action for an AI agent?

That is one of the clearest reported use cases. Keep the action set bounded. Validate the selected tool, arguments, permissions, resource scope, and confirmation requirement before execution.

Is a confidence score the same as accuracy?

No. Confidence may help rank or route cases, but it must be evaluated against real outcomes. For costly decisions, add thresholds, abstention, human review, deterministic rules, and monitoring.

Should developers use it now?

Use it for a controlled experiment if your account has access and your workflow can tolerate preview changes. Start with offline or shadow evaluation, not irreversible automation. Study the typed-decision pattern before designing your adapter.

What is the main takeaway?

The Decisions API represents a shift from “ask a model to write something” toward “ask a model to make one small, typed judgment.” That can make classification, routing, and agent control loops faster and easier to integrate. The engineering discipline remains the same: define the answer space, evaluate on representative data, keep authority in code, and treat confidence as evidence rather than permission.

Sources

  • OpenAI DevDay 2026 announcement
  • The Decoder: OpenAI expands Codex and its API at DevDay.
  • Pasquale Pillitteri: OpenAI launches Decisions API to take on Jev.
© 2026 Jev AI JournalBack home