OpenAI Decisions API Tutorial [2026]: 3 Cost-Safe Patterns
A production-first OpenAI Decisions API tutorial: typed evals on POST /v1/decisions, predicate/choice/score, and three patterns to gate, batch, and escalate without cost blowups.
If you want a fast, typed “yes/no / pick-one / score-it” evaluator in production, the OpenAI Decisions API is the cleanest thing OpenAI has shipped in a while.
This openai decisions api tutorial walks through the mental model, a minimal integration, and three patterns I’d actually ship: gate → generate, cascade (rules → Decisions → Responses), and batch + cache.
Here’s the part that trips people up. Decisions is not a chat endpoint. You’re not “asking the model to respond.” You’re sending shared evidence (input) plus multiple typed questions, and you’re getting back typed answers with probabilities.
What is OpenAI Decisions API
OpenAI Decisions API is a dedicated endpoint that evaluates text, images, or both and returns typed answers (predicate/choice/score) about 10x faster than the Responses API, according to OpenAI’s docs.

It solves a specific production problem: you need a model to make a decision (classification, routing, prioritization, triage) quickly and consistently, without paying for long-form generation you’re going to throw away anyway.
OpenAI says the Decisions API is in public beta, they “expect to GA in the coming weeks,” and `gpt-6-luna` is currently the only supported model. Source: OpenAI API docs.
How decisions work
Here’s the mental model I use.

- You send one shared `input`. That’s the evidence.
- You send many `questions`. Each question has a
type,name, and instructions. - You get back an `answers` array, keyed by
name.
OpenAI exposes Decisions on a dedicated endpoint: `POST /v1/decisions`. That separation matters operationally. I treat it like a router service. Low-latency, high-volume, cheap-ish calls that decide whether we ever wake up expensive generation.
This is also where I see teams light money on fire.
They call a big generative model to “think.” Then they call a classifier model to “decide.” Then they call the big model again to “write.” Decisions is meant to replace that middle step and, in a bunch of cases, remove the first one too.
One lesson from running this site’s multi-agent publishing pipeline (261+ posts so far): deterministic gates beat bigger models. Same philosophy here. Use Decisions outputs as a gate before you spin up a long, expensive model session.
Minimal Decisions request (one input, multiple questions)
Below is a minimal, runnable JavaScript example using the OpenAI SDK. The whole point is that you pack multiple checks into one call.
import OpenAI from "openai";
const client = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });
const res = await client.decisions.create({
model: "gpt-6-luna",
input: [
{
role: "user",
content: "Customer message: 'My package arrived cracked. The screen is shattered. I need a refund.'",
},
],
questions: [
{
name: "is_refund_request",
type: "predicate",
instructions: "Is the customer explicitly requesting a refund or return?",
},
{
name: "issue_category",
type: "choice",
instructions: "Classify the primary issue.",
choices: ["shipping_damage", "billing", "setup_help", "other"],
},
{
name: "severity",
type: "score",
instructions: "Rate severity for support prioritization.",
levels: [
{ label: "low", description: "Minor inconvenience, no device failure" },
{ label: "medium", description: "Device usable but impaired" },
{ label: "high", description: "Device unusable or safety risk" },
],
},
],
});
for (const ans of res.answers) {
console.log(ans.name, ans);
}Notes that actually matter once this is in production:
- The shared `input` is amortized across all questions. If you have 6 predicates, do not do 6 calls.
- Give each question a stable
name. That becomes your metric key, your dashboard label, and your alert dimension. - Treat
probabilitylike a signal you have to validate, not a truth machine.
Choose a question type
Decisions supports three question types: `predicate`, `choice`, and `score`. OpenAI’s guide calls out the core semantics, and the differences matter if you want to wire outputs into real logic without embarrassing yourself later: OpenAI API docs.

`predicate`: probability that a condition is true
You get probability in [0, 1]. This is your workhorse for gating.
How I map this in production:
- High confidence (
p ≥ 0.9): auto-route with no escalation. - Uncertainty band (
0.4 ≤ p ≤ 0.6): escalate to a more expensive call. - Low confidence (
p ≤ 0.1): treat as false and move on.
The exact thresholds depend on what you’re optimizing for. But the uncertainty band pattern is non-negotiable.
`choice`: one of N labels, plus probabilities
Use this when labels aren’t ordered. Classic example: department routing.
How I map this:
- Use
choicedirectly for routing. - Use the top probability as your confidence score.
- If top probability is low (say
< 0.7), fall back to human review or a generative model with Structured Outputs.
`score`: ordered levels with a probability-weighted average
This is the one people mess up.
In OpenAI’s definition, score is the probability-weighted average of level indices. So if the model is 50/50 between “medium” and “high,” your score might land around 1.5 (if indices are 1 and 2). That’s not a bug. That’s the model telling you it’s torn.
How I map this:
- Don’t round blindly. Use bands.
- Scores are great for queues: sort descending, then cap how many get escalated per minute/hour/day.
Rubric design rule I follow
I keep score levels to 3–5. Fewer than 3 is too coarse. More than 5 is fake precision dressed up as engineering.
Write level descriptions like a runbook. If an on-call engineer can’t apply the rubric at 2am, the model won’t either.
Check an image for visible damage
OpenAI’s docs include an example using a predicate to check an image for visible damage. The important trick is you can mix text instructions with an image in the shared input.
Here’s a minimal version in JavaScript. (Use any publicly accessible image URL or a file upload flow you already have.)
import OpenAI from "openai";
const client = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });
const imageUrl = "https://example.com/product-photo.jpg";
const res = await client.decisions.create({
model: "gpt-6-luna",
input: [
{
role: "user",
content: [
{ type: "input_text", text: "Check the photo for visible damage." },
{ type: "input_image", image_url: imageUrl },
],
},
],
questions: [
{
name: "visible_damage",
type: "predicate",
instructions:
"Is there visible damage to the item (cracks, dents, broken parts) in the photo?",
},
],
});
const damage = res.answers.find((a) => a.name === "visible_damage");
console.log(damage);In a real system, I’d pair this with a second predicate like “is the photo usable (not too blurry/dark).” Otherwise you end up pretending you made a reliable call from garbage input, which is how bad automations get shipped.
Decisions vs Structured Outputs vs function calling (when to use what)
This is where most stacks go sideways.
- Decisions is for typed evaluation.
- Structured Outputs is for typed extraction / generation.
- Function calling is for actions (tools), often inside AI agents.
OpenAI’s positioning is pretty clear: use Structured Outputs with the Responses API when you need JSON that matches your schema, and use function calling when you want the model to request a tool call with arguments (OpenAI API docs and OpenAI API docs).
Here’s the decision table I use:
| Need | Use | Why | Output shape |
|---|---|---|---|
| Fast classification / routing / triage | Decisions API | Cheap(ish), typed, probability-bearing | predicate/choice/score + probabilities |
| Extract fields into your domain model | Responses + Structured Outputs | You need your schema, not a rubric | JSON object (schema-valid) |
| Take an action (call a function/API) | Function calling / Agents | You need side effects + arguments | tool call request + args |
How Decisions differs from tool-calling and agents, in one line: agents decide and then act. Decisions only decides.
If you’re building AI agents, Decisions is a great front door router so the agent doesn’t wake up for junk requests.
Also, Decisions can reduce your attack surface. If your agent uses web search or tools, put a Decisions gate in front of it to classify intent and risk. I wrote about this failure mode in prompt injection scenarios, where an agent ends up executing tool calls because it never had a strong “should I even do this?” gate.
3 integration patterns that don’t blow up costs
If you do nothing else, do these three things.
Pattern 1: Cheap gate → expensive generate (with an uncertainty band)
This is the “don’t page the expensive model unless it matters” pattern.
Algorithm:
- Call Decisions
predicatefor a key condition. - If
p ≥ high_threshold, skip generation and route automatically. - If
p ≤ low_threshold, skip generation and route the other way. - If
low_threshold < p < high_threshold, escalate to Responses.
Concrete numbers (start here, then tune):
high_threshold = 0.85low_threshold = 0.15
That’s a wide uncertainty band on purpose. The point is not being “smart.” The point is controlling your escalation rate.
type GateDecision =
| { action: "auto_route"; reason: string }
| { action: "escalate"; reason: string };
export function gateWithUncertaintyBand(p: number): GateDecision {
if (p >= 0.85) return { action: "auto_route", reason: `high_confidence:${p}` };
if (p <= 0.15) return { action: "auto_route", reason: `low_confidence:${p}` };
return { action: "escalate", reason: `uncertain:${p}` };
}Cost control math you should actually do:
If Decisions is ~10x faster than Responses (per OpenAI), teams will immediately slap it on every request. Fine, until your volume is 1M/day.
So estimate it. Don’t hand-wave it.
- Volume: 1,000,000 requests/day
- Escalation rate after gating: 5% (50,000/day)
Even if your Decisions spend doubles, cutting Responses calls by 95% changes your unit economics completely.
If your escalation rate is 40%, you didn’t build a gate. You built an extra step.
If you want a deeper way to do this math across retries, caching, and tool calls, see AI in production and my breakdown of LLM cost.
Pattern 2: Cascade (rules → Decisions → Responses)
This is one of those things where the boring answer is actually the right one.
Before you call any model, run dumb rules.
- If the request is empty, drop it.
- If the user is on a free tier, cap features.
- If the request matches a known safe template, route it.
Then call Decisions.
Then call Responses only for the uncertain band.
def route_request(req):
# Stage 0: deterministic rules
if req.user_plan == "free" and req.kind == "bulk":
return {"action": "reject", "reason": "plan_limit"}
if len(req.text) < 10:
return {"action": "ask_clarifying", "reason": "too_short"}
# Stage 1: Decisions
p = decisions_predicate_probability(req.text, "is_high_risk")
if p >= 0.9:
return {"action": "send_to_human", "reason": f"risk:{p}"}
if p <= 0.2:
return {"action": "auto_handle", "reason": f"risk:{p}"}
# Stage 2: expensive generation
return responses_structured(req.text)Why this works: the cascade turns your “LLM system” into a budgeted pipeline.
I do the same thing in this blog’s agent pipeline. The deterministic SEO quality gate runs before the expensive review model because catching issues early is cheaper than trying to argue with a bigger model later.
Pattern 3: Pack-many-questions + cache
OpenAI designed Decisions around one shared input, many questions. Use it.
Batching strategies:
- Pack 5–20 predicates into one call when they share the same evidence.
- Prefer one
choicewith 6 categories over 6 predicates that fight each other. - Add one
scorefor priority instead of multiple severity predicates.
Caching strategies:
- Cache Decisions results keyed by a stable hash of
(normalized_input, questions_version). - TTL based on risk. Example: “toxicity” cache for 7 days. “is this invoice paid?” cache for 5 minutes.
Even a modest cache hit rate matters. If you hit 30% cache on a high-volume endpoint, you just bought yourself headroom for future product growth.
If you’re already using semantic caching for generation, don’t reuse it for Decisions. The failure modes are different. For Decisions, you want strict equality and versioning.
Production mapping: from probabilities to actions (without fooling yourself)
This is where typed outputs earn their keep.
Thresholds: pick them backwards from budgets
Start with what you can afford.
Example:
- You can afford to escalate ≤ 8% of requests to Responses.
- So you set thresholds to produce ~8% uncertainty-band outcomes on your traffic.
Then you evaluate accuracy on that slice.
This is the part nobody wants to hear: your thresholds are a product decision. Not a model decision.
Safest defaults when the Decisions call fails
Failures happen. Timeouts, 5xx, transient network issues, the usual.
Safe defaults depend on blast radius:
- If the decision gates money movement (refunds, credits): default to escalate to human.
- If the decision gates UX quality (which template to show): default to cheaper deterministic path.
- If the decision gates an agent/tool call: default to deny tool use.
This pairs nicely with spend limits. If you haven’t implemented them yet, read AI agent kill switch spend limits.
What to log and monitor in production
If you ship Decisions without monitoring, you’re basically blind.
Here’s what I log per request:
decisions_model:gpt-6-luna(today) and whatever replaces it laterquestion_nameanswer_type: predicate/choice/scoreprobability(predicate) or top-1 probability (choice)score(score)latency_msescalated: booleanfallback_reason: timeout/5xx/parse/etc.
And here’s what I monitor:
- Probability distribution drift per question (histograms by day). If your
is_spamprobabilities shift from being bimodal to clustered around 0.5, something changed. - Escalation rate. If it climbs from 8% to 22%, your costs are about to follow.
- Error rate and timeout rate. If Decisions starts timing out, your uncertainty band logic will accidentally stampede Responses unless you cap escalation.
For a broader monitoring schema (including redaction and trace IDs), see AI in production and my template for agent orchestration.
My prediction
Decisions is going to make a lot of agent stacks look silly.
Not because agents are bad. Because half of what people call “agents” is really just routing and triage wrapped in a chat loop.
Typed evaluators are going to peel that logic out into a cheaper, faster front door. And once teams see the bills stop spiking, they won’t go back.
If you’re building an agent system this quarter, my challenge is simple: put Decisions in front of it and prove, with numbers, that your escalation rate stays under 10%. If you can’t, you’re not building an agent. You’re building a cost leak with a nice demo.
Photo by Levart_Photographer on Unsplash.
Kunal Ganglani (2026, October 7). OpenAI Decisions API Tutorial [2026]: 3 Cost-Safe Patterns. Kunal Ganglani. Retrieved October 7, 2026, from https://www.kunalganglani.com/blog/openai-decisions-api-tutorial



