Skip to main content Announcing Tool Gateway MCP: the universal MCPRead the announcement
Guillaume Lebedel Guillaume Lebedel · · 6 min
OpenAI Decisions API request with three question types, predicate, choice and score, returning typed answers

What is OpenAI's Decisions API, and when to use it

Table of Contents

OpenAI’s Decisions API is a public beta endpoint that runs GPT-6 Luna on bounded questions: the probability a condition is true, one choice from a fixed list, or a score against ordered levels. OpenAI says it answers about 10x faster than the Responses API and bills only input tokens, at $0.10 per million. It does not write tool arguments.

Facts checked against vendor docs and public posts on 7 October 2026. The API is in beta, so limits and pricing may change.

What can OpenAI’s Decisions API decide?

It answers three kinds of question about text, images or both. A predicate returns the probability that a condition holds. A choice returns one option from a list you define, with a probability for each option and a confidence value. A score rates the input against ordered levels, such as ticket severity, and returns a probability-weighted position on that scale.

You send model, input and a questions array to POST /v1/decisions, and get back an answers array with one typed answer per named question. According to the OpenAI Decisions guide, only gpt-6-luna is supported for now and images must be base64 inline. The public beta announcement went out on 6 October 2026, a week after OpenAI first showed the API at DevDay.

How is OpenAI’s Decisions API different from function calling and Structured Outputs?

The Decisions API picks from answers you wrote. Structured Outputs generates an object that fits your JSON schema, and function calling generates a tool name plus its arguments. OpenAI’s own guide draws the boundary: use Structured Outputs for extracted fields or written explanations, “or function calling when you need a model to request a tool call with arguments.”

Decisions APIStructured OutputsFunction calling
What comes backA probability, a choice or a scoreAn object matching your schemaA tool name and generated arguments
Generates textNoYesYes, the arguments
Billed tokensInput only, $0.10 per 1MInput and outputInput and output
Typical jobClassify, gate, routeExtract fields, explainCall a tool with parameters

A routing call costs the same whatever answer comes back, because there is no output to bill. Budget on cost per completed task rather than price per token, though: a decider that hands half its cases to a large model is not cheap.

How does OpenAI’s Decisions API compare with Jev on cost and accuracy?

On list price, OpenAI is about 2.4 times more expensive than TypeSafe AI’s Jev, a model built only for decisions: $0.10 against $0.042 per million input tokens, with no output charge on either. The first outside benchmark, published on 7 October 2026, found OpenAI 2x more expensive and 5 to 10% less accurate on a production task.

Bar chart of input price per million tokens: OpenAI Decisions API with gpt-6-luna at $0.10 and TypeSafe AI Jev 1.13 at $0.042, both with no output charge

That benchmark comes from Hamed Nilforoshan’s test for HiringCafe, a job search site serving 2.5 million users. The task scored how relevant a query or resume is to a job description on a 1 to 10 scale, which maps to a score question. On a few hundred word-game and conversation questions, a developer on the OpenAI community forum reported something similar: Luna cost about 2 to 3 times as much and made about three times as many confident wrong answers on nuanced judgment calls, while narrow yes/no checks came out about even.

Both are single tests from the first day of the beta, so treat them as an early signal. Jev’s price comes from its OpenRouter listing, live since 18 September 2026.

OpenAI’s clear advantage is image input. Jev takes a text or JSON state, so a decision that depends on a screenshot needs OpenAI’s endpoint for now.

Where does a decision model fit in an agent loop?

A decision model sits in front of the expensive model as a fast, cheap router. It classifies the request, gates risky steps and picks a path, and your code acts on the typed answer. Anything that needs generated text or tool arguments goes on to a general model.

Diagram of an agent loop: input goes to a decision model that classifies, gates and routes; simple answers branch in code, while hard cases go to an LLM that writes tool arguments, which are validated before the action runs in the business app

Typical jobs for the decider in an agent built for HR, CRM or IT service management:

  • Route an inbound request to the right team or workflow.
  • Gate a write action: is this change within the task the user asked for?
  • Score urgency or severity before deciding whether a human reviews it.

The argument step is where agents most often break: right tool, wrong field, or a value the target system rejects. A decider does not fix that, because it never writes the arguments.

Can a decision model choose an agent’s tools and actions?

It can choose which tool to call from a short list, and OpenAI’s announcement names choosing “the right model, tool, or action” as a use case. Writing that tool’s arguments still needs function calling, so for actions against business apps the decider covers half the job.

In our test of yes/no tool selection across 14 models this summer, asking a model a plain yes or no per tool improved selection accuracy by 14.9 points on average, but the effect swung widely by model and did not cover parameter extraction.

That split between choosing an action and filling it in is why our team built StackOne’s tool search the way it is. The agent gets two tools: one searches a catalogue of 33,000+ actions across 540+ connectors in natural language, and one executes the action it picked with the parameters the model writes. Loading every tool definition up front is one of the bigger places agent tokens go, and with search the prompt stays the same size however many apps are connected. A decision model could rank those search results too; we have not measured that yet.

How should you test a decision model before switching to it?

Build a labelled set from your own traffic and run both models on it before you switch. Launch benchmarks rarely match your task, and the two early tests above show results can move by several points of accuracy and 2 to 3 times on cost.

  1. Pull a few hundred real cases per decision type from production logs and label the correct answer.
  2. Run each candidate on identical inputs with identical question definitions.
  3. Shuffle the option order on choice questions and rerun. TypeSafe’s Jev failure-mode guide lists option order as a known issue, and a forum tester saw the same effect in OpenAI’s beta.
  4. At the confidence threshold you plan to act on, count the confident wrong answers, not only overall accuracy.
  5. Compute cost per completed decision, including the cases you hand on to a larger model.
  6. Rerun the set on every new model version.

The full method is in how to re-benchmark an agent before a model swap.

OpenAI may well close the gap with Jev quickly. A labelled set of your own cases tells you when it has.

To see how search and execute works in an agent, the Advanced Tool Search docs walk through the two-tool setup and the agent frameworks it supports.

Frequently Asked Questions

Is the Decisions API generally available?
The Decisions API is not generally available yet: it has been in public beta for all developers since 6 October 2026, after a limited DevDay preview. OpenAI expects general availability in the coming weeks, so endpoint behaviour, limits and pricing can still change.
Does the Decisions API charge for output tokens?
The Decisions API does not charge for output tokens. With gpt-6-luna, OpenAI bills $0.10 per million input tokens and nothing for output, cache reads or cache writes. Regional processing and long-context requests carry multipliers. Cost per call therefore depends on how much context you send.
Does the Decisions API support zero data retention?
The Decisions API supports zero data retention for qualifying customers. OpenAI's guide also lists HIPAA eligibility and regional processing in the US and Europe for this endpoint. Confirm your organisation qualifies before you send personal or health data through a beta API.
Which Decisions API question type should you start with?
The easiest Decisions API question type to start with is the predicate, because a single probability is the easiest output to threshold and audit. Use choice for routing between a handful of named paths, and score for ordered ratings such as severity. In the beta, a forum tester found predicate answers tracked expected probabilities more closely than choice answers.

Put your AI agents to work

All the tools you need to build and scale AI agent integrations, with best-in-class connectivity, execution, and security.