News

OpenAI launches the Decisions API: typed AI answers, and output costs nothing

OpenAI's new Decisions API returns probabilities, choices and scores instead of text, bills only input tokens and claims ten times the speed. Here is when it saves money and when it does not.

OpenAI Decisions API news card

OpenAI has opened a new kind of endpoint for developers, and it breaks with one of the oldest habits in AI pricing. The Decisions API went into public beta on 6 October 2026. It does not write essays, code or emails. It answers questions with a number, a pick from a list or a score on a scale. And it bills only what you send in: there are no charges for output tokens, and no charges for cache reads or cache writes either.

Input costs 0.10 US dollars per million tokens on GPT 6 Luna, the only model the endpoint supports for now. OpenAI says the answers come back about ten times faster than the same work through its general purpose Responses API. General availability is expected "in the coming weeks", according to the official Decisions guide.

This article explains what the Decisions API is, what the three question types do, what it really costs once you run the numbers, where it is cheaper and where it is surprisingly more expensive, and what it tells us about where the AI market is heading. Pricing and speed claims come from OpenAI itself. Additional context comes from reporting by The Decoder and from analyses published after the launch. The cost scenarios further down are our own calculations based on OpenAI's published rates.

Why an API that only says yes or no

Most people think of large language models as text machines. You ask, they write. But a huge share of real world AI traffic is not about writing at all. Companies use models to sort support tickets, flag harmful posts, decide whether an invoice matches an order, check whether a product photo shows damage, rank search results or pick which team should handle a request. In all of these jobs the model reads a lot and says very little.

Until now, developers solved this by asking a chat model to "answer only with yes or no" or by forcing a small JSON object through structured outputs. That works, but it is clumsy. The model still generates tokens one by one, you pay for every one of them, and you have to parse the reply and hope it follows the format. Worse, a plain text "yes" tells you nothing about how sure the model is.

The Decisions API turns this pattern into a product. You hand over the evidence, you list the questions, and you get back typed answers with probabilities attached. There is nothing to parse and nothing to clean up.

How a request works

A Decisions request has three parts. The model field names the model, today always `gpt-6-luna`. The input carries the shared evidence: a text string, or user messages that mix text and images. The questions array lists what should be evaluated, each with a unique name, a type and instructions. The response contains an answers array, and each answer echoes the name of its question, so your code always knows which result belongs to which question.

You can ask several independent questions about the same input in one call. A shop could check a product photo for visible damage and classify the product category in a single request. When one decision depends on another, OpenAI recommends separate requests: first check for damage, then ask for the repair category only if damage was found.

Images have one notable limit. They must be sent inline as base64 data URLs. Hosted image links and uploaded file IDs are not supported on this endpoint, so an app that stores pictures in the cloud has to fetch and encode them first.

Overview of the three Decisions API question types: predicate returns a probability, choice picks one value from a list, score rates against ordered levels

The three question types

Predicate questions check whether a condition is true. The answer is a probability between 0 and 1. In OpenAI's own example, the model inspects a product photo and estimates how likely it is that the item has a crack, tear or dent, while ignoring shadows and damage to the packaging. Your application then decides what threshold triggers a human review.

Choice questions select one value from a fixed list that you provide. The classic case is routing: a complaint such as "I was charged twice for my order" should land in billing, not in technical support or shipping. The answer contains the chosen value, a probability for every option and a separate confidence figure. OpenAI advises adding a fallback option such as "other", so inputs that fit no category can go to a general queue instead of being forced into the wrong box.

Score questions rate an input against ordered levels, for example cosmetic, workaround available and fully blocked for a software bug. The score is a probability weighted average of the level positions, so it can land between two levels. If the model is 10 percent sure a bug is cosmetic, 70 percent sure there is a workaround and 20 percent sure the user is fully blocked, the score comes out at 1.1. That is more honest than a single forced label, because it shows the model's uncertainty.

The advice in the guide is practical: write questions around observable criteria, keep separate concerns in separate questions, give choices distinct meanings and make sure neighbouring score levels are clearly different. Then calibrate thresholds with labelled examples from your own data, weighing the cost of a false alarm against the cost of a miss.

What it costs

The headline price is simple. On the Decisions endpoint, GPT 6 Luna input costs 0.10 dollars per million tokens. Output, cache reads and cache writes cost nothing. Two surcharges still apply, according to the guide and to analyses of the pricing page: regional processing in the United States or Europe adds a premium, reported at 10 percent, and long requests above 272,000 input tokens are billed at a higher rate for the whole request.

For comparison, GPT 6 Luna through the normal Responses API costs the same 0.10 dollars per million input tokens, but only 0.01 dollars for cached input, and 0.50 dollars per million output tokens.

Bar chart of GPT 6 Luna token prices: input 10 cents per million on both routes, cached input 1 cent on Responses, output 50 cents on Responses and 0 cents on Decisions

So the free output is real, but it is not as dramatic as it sounds. Classification answers are short anyway. A routing reply of 50 tokens on the Responses route costs very little. Let us run a typical example.

Imagine 10,000 support messages of about 1,500 tokens each that need to be routed to the right team. That is 15 million input tokens. On the Decisions API the bill is 1.50 dollars. On the Responses API the same input also costs 1.50 dollars, plus 500,000 output tokens for 50 token answers, which adds 0.25 dollars. Total: 1.75 dollars. The Decisions API saves around 14 percent in this case, before you count the speed gain and the simpler code.

Bar chart of the cost for 10,000 routing decisions of 1,500 tokens: 1.50 dollars on the Decisions API versus 1.75 dollars on the Responses API

The catch: no cache discount

There is one scenario where the new endpoint gets more expensive, and it is common. Many companies send the same long instruction block with every call: a detailed policy, a labelling rubric or a list of examples. On the Responses API, that repeated prefix is cached and billed at a tenth of the normal input rate. On the Decisions API there is no separate cache rate. Every token counts at the full 0.10 dollars.

Take a 2,000 token rubric that stays the same for every call, plus a 200 token item that changes. On the Responses API with caching, 1,000 calls cost roughly 5 cents: about 2 cents for the cached rubric, 2 cents for the new items and 1 cent for short answers. On the Decisions API the same 1,000 calls cost about 22 cents, because the whole 2,200 tokens are billed at the full rate. That is more than four times as much.

Bar chart of the cost per 1,000 calls with a 2,000 token cached rubric: about 22 cents on the Decisions API versus about 5 cents on the Responses API with caching

The lesson is simple. The Decisions API is a great deal for many short, varied items without a long shared preamble. If your workload repeats a big prompt thousands of times, do the maths before you switch. In absolute terms both numbers are tiny, but at hundreds of millions of calls the gap becomes real money.

Speed matters more than price

For many teams the ten times speed claim will be the bigger story. A moderation filter that must decide before a post goes live, a voice assistant that picks an action while the user is still talking, or a fraud check in a checkout flow all live or die by latency. Because the model does not have to generate a reply token by token, a typed answer can come back much faster.

OpenAI also links the endpoint to its Live API. Through what it calls client delegation, a voice app can let the Decisions API choose actions from spoken requests and report the results back to the user. That points to a future where small, fast decision calls sit between a conversation and the tools behind it.

Speed figures are OpenAI's own. Independent benchmarks were not available at launch, and real latency will depend on input size, images and region.

Privacy and compliance

The endpoint supports Zero Data Retention and HIPAA use for eligible customers, according to the guide. Data residency and regional processing are available in the United States and in Europe, covering the EEA and Switzerland. For European companies that want to classify customer messages or documents without sending data out of the region, that is an important detail, although the regional option comes with its price premium.

Simpler API tiers on the same day

The same changelog entry also reshaped OpenAI's usage tiers for the API. According to reports on the changelog, the five paid tiers become three, called Build, Launch and Grow, with automatic upgrades as total credit purchases reach each minimum. Build is reported to start at 5 dollars purchased with a 500 dollar monthly usage limit, Launch at 100 dollars with 5,000 dollars a month and Grow at 500 dollars with up to 200,000 dollars a month. For small developers, that means fewer steps before they reach usable rate limits.

How it fits the price war

The Decisions API arrives in the middle of an aggressive price fight between the big model labs. Anthropic's Claude Haiku 5.5 launched this week at 0.10 dollars per million input tokens and 0.50 dollars per million output tokens, the same list price as GPT 6 Luna. Google halved the price of its image model the same week. Cheap small models are turning into a commodity, so the labs compete on packaging: faster endpoints, typed outputs, fewer tokens to pay for.

Removing the output charge for one narrow job is a clever move in that fight. It makes the price easy to explain, it rewards developers for using the most efficient format, and it pulls classification and routing workloads, which are huge in volume, onto OpenAI's infrastructure. Once a company's moderation or routing pipeline runs on an endpoint like this, it rarely moves.

Who should try it

The Decisions API is a strong fit if you:

  • classify, route or grade many short, varied inputs such as tickets, posts, emails or photos
  • need a confidence value to decide when a human should look
  • care about speed, for example in moderation, checkout or voice apps
  • want to stop writing fragile parsing code for yes or no answers

It is a weaker fit if you:

  • reuse a very long, identical rubric or set of examples in every call
  • need the model to explain its reasoning or extract fields into your own JSON schema, where OpenAI itself points to structured outputs
  • need models other than GPT 6 Luna
  • store images only as hosted links and do not want to encode them

What to watch

Betas change. Prices, limits and supported models can move before general availability, and OpenAI's main pricing page did not yet list a separate Decisions row in the days after launch, according to one analysis. The two numbers that make this endpoint interesting, free output and ten times the speed, are both vendor claims. The real test is your own traffic.

Still, the idea is bigger than one endpoint. AI is increasingly used not to talk, but to decide. An API built around that, with probabilities instead of prose and with a price that charges only for reading, is a sign of how the business is maturing. Expect competitors to follow.

Sources

Source: developers.openai.com

Newsletter

The AI news that matters, in your inbox.