OpenAI / News

OpenAI's Decisions API: ten times faster, but watch the order

OpenAI's Decisions API answered our 60 support messages in about 125 milliseconds, ten times faster than the normal way. But changing the order of the options changed 6 answers.

Bar chart of our test: the Decisions API answered in 125 milliseconds in the middle, the Responses API in 1210 milliseconds with thinking off and 1696 milliseconds with default thinking
Our own test with GPT-6 Luna, 8 October 2026. Chart by Not an AI App.

OpenAI has a new tool for developers: the Decisions API. You give it a text or a picture and a few fixed questions, and it gives back short answers with numbers. OpenAI says it is about ten times faster than asking the same model in the normal way. So we tested that. In our test with 60 support messages, it answered in about 125 milliseconds. The normal way took 1.2 to 1.7 seconds. But when we only changed the order of the answer options, 6 of the 60 answers changed too.

What is the Decisions API?

Screenshot of OpenAI's guide: the Decisions API evaluates text, images, or both and returns typed answers about 10x faster than the Responses API
From OpenAI's Decisions API guide. Screenshot from 8 October 2026, cropped.

An API is the connection programs use to talk to an AI. Most AI APIs let the model write a full answer, word by word. The Decisions API does not. It only fills in the answers to questions you set up yourself. OpenAI calls them "typed answers": a number, or one of the options you gave.

That is useful for jobs a program does thousands of times a day. For example: sorting customer emails, checking if a photo shows damage, or choosing what an AI helper should do next. OpenAI launched it on 6 October 2026 as a public beta, which is a test version that everyone may use. It only works with one model, GPT-6 Luna, OpenAI says in its guide.

Three kinds of questions

Screenshot of OpenAI's table of question types: predicate gives a probability from 0 to 1, choice gives one of your values, score gives a weighted average of ordered levels
The three question types in OpenAI's guide. Screenshot from 8 October 2026, cropped.

You can ask three kinds of questions:

  • Predicate: a yes-or-no question. You get back a chance from 0 to 1. For example: "Does the customer ask for money back?" 0.95 means "almost certainly yes".
  • Choice: pick one of your options. For example: billing, technical, account, shipping or other. You also get a chance for each option and a "confidence" number.
  • Score: a rating on a scale you describe, like low, medium or high. The score can land between two levels, like 1.3.

You can put several questions in one request. They all look at the same text or picture.

What you send and what you get back

Our real request with a support message and three questions, and the answer: category billing with confidence 1, refund probability 1, urgency score 1.29
A real request we sent on 8 October 2026, shortened. Layout by Not an AI App.

Here is a real request we sent. The message was: "I was charged twice for my order #4411 last week. Please send the extra money back." We asked three questions: which team, does the customer want a refund, and how urgent is it.

The answer came back as numbers, not as a story. Team: billing, with confidence 1. Refund: a chance of 1. Urgency: 1.29, so a bit above "medium". The API read 430 tokens and wrote none.

One limit to know: pictures must be sent inside the request itself. A link to a picture on a website does not work, the guide says.

What does it cost?

Screenshot of OpenAI's guide: with gpt-6-luna, input costs $0.10 per 1M tokens, and you pay only for input tokens
From OpenAI's Decisions API guide. Screenshot from 8 October 2026, cropped.

You only pay for what the API reads: $0.10 per million tokens. There is no charge for the answer, OpenAI says. That is the same price per token as reading with GPT-6 Luna in the normal way.

But there is a catch. Your questions and options also count as text to read, and they are sent again with every message. In our test, the Decisions API read about 361 tokens per message. The normal way read about 187 and wrote a short answer. So in the end, sorting 1,000 messages cost:

  • Decisions API: $0.036
  • Normal way, thinking off: $0.027
  • Normal way, default thinking: $0.038

All three are very cheap. But the Decisions API is not cheaper per message. It is faster.

Our test: 60 support messages

Screenshot of OpenAI's guide: use labeled examples from your application to set thresholds for routing, filtering, or review
OpenAI's own advice in the guide. Screenshot from 8 October 2026, cropped.

OpenAI itself advises testing with your own labelled examples. So we wrote 60 short support messages, like a web shop could get. For each one we decided which team should handle it, and whether the customer asks for money back. Some examples:

  • "Photos won't upload. I get error 503 after the progress bar reaches 90%." → technical, no refund.
  • "The box arrived crushed and the mug inside is broken. I want my money back." → shipping, refund.
  • "Close my account today and refund the rest of my yearly subscription." → account, refund. This one is tricky: billing would also make sense.

We sent each message to GPT-6 Luna in four ways:

  • the Decisions API, with the teams in our order;
  • the Decisions API, with the teams in reverse order;
  • the normal Responses API, with thinking off;
  • the normal Responses API, with the default thinking.

We sent one message at a time, so the times did not get in each other's way.

Want to try it yourself? Download our test kit: the 60 messages with our labels and the script we used. It is free to use under the MIT licence. You need your own OpenAI API key; one run costs about one cent.

Is it really ten times faster?

Screenshot of OpenAI's changelog for 6 October: Released the Decisions API in beta with gpt-6-luna, 10x faster than the Responses API
OpenAI's developer changelog. Screenshot from 8 October 2026, cropped.

Yes, in our test it was. The middle time for one answer was:

  • Decisions API: 125 milliseconds
  • Responses API, thinking off: 1210 milliseconds
  • Responses API, default thinking: 1696 milliseconds

So the Decisions API was about 10 to 14 times faster. 1,000 milliseconds is one second. For one email that difference does not matter much. For an app that must decide something while you wait, or for a voice assistant, it matters a lot.

We measured from our computer in the Netherlands. The very first request took almost two seconds; after that it was fast.

How often was it right?

Table of our test: Decisions API 57 of 60 right teams, 55 with the teams in reverse order, Responses API 59 and 58; refund question 60 of 60 for all
Our own test, 8 October 2026. Costs worked out from the token counts the API returned.

On the refund question, every way got all 60 right.

On the team question, the Decisions API chose the team we expected for 57 of the 60 messages. A second run gave exactly the same answers. The Responses API got 58 with default thinking and 59 with thinking off. When we ran the kit again later with thinking off, it got 60 of 60. So small differences like this can be luck.

The misses were mostly messages that fit two teams. For example: "Your software has been broken for two weeks… I want my money back for this month." We said technical, but both APIs said billing. That is not really wrong.

Change the order, change the answer

Table of the 6 messages that got another team when we only reversed the order of the five teams in the question
Our own test, 8 October 2026. Chart by Not an AI App.

This was the biggest surprise. We sent the same 60 messages again, but with the five teams listed in reverse order. Nothing else changed. Still, 6 of the 60 answers changed. The Decisions API now got 55 right instead of 57.

For example, "Do you offer next-day delivery?" first went to shipping, and then to "other". And "Can I add my colleague to my account?" went from account to "other".

We were not the only ones to see this. On OpenAI's forum, developers who tried it in the first days also reported that the order of the options changed the chances. The yes-or-no answers (predicates) did not have this problem in our test: the refund answers stayed the same.

Can you trust the confidence number?

Bar chart: all 42 answers with a confidence of 0.9 or more were right; lower numbers held both right and wrong answers
Our own test, 8 October 2026, teams in our order. Chart by Not an AI App.

With every choice, the API gives a confidence number from 0 to 1. Can you use it to spot the doubtful answers?

Partly. All 42 answers with a confidence of 0.9 or more were right. The 3 wrong answers had a confidence between 0.41 and 0.77. But 7 right answers also had a confidence below 0.7.

So a simple rule could be: trust answers of 0.9 or more, and let a person check the rest. In our test, that would have sent 18 of the 60 messages to a person, and caught all the mistakes. With your own messages, the best line may be somewhere else. OpenAI also says you should pick that line yourself, with your own examples.

Who is it for?

Screenshot of OpenAI's forum post: the Decisions API lets your app choose the right model, tool or action in near real-time, and is up to 10x faster than GPT-6 Luna through the Responses API
OpenAI's announcement on its developer forum, 6 October 2026. Screenshot from 8 October 2026, cropped.

The Decisions API is a tool for developers, not for people who use ChatGPT. It makes most sense for apps that must make many small decisions very quickly. Think of sorting messages, checking photos or choosing the next step for an AI helper.

A few things to know before you use it:

  • It is still a beta. OpenAI expects the final version "in the coming weeks".
  • It only works with GPT-6 Luna.
  • For companies that qualify, OpenAI offers zero data retention, which means OpenAI does not keep your data, and data storage in the US or Europe.
  • Use yes-or-no questions where you can, and test choices in more than one order.

GPT-6 Luna also took part in our coding test for Claude Haiku 5.5, where it did well for a very low price.

How we checked

Diagram of our method in three steps: read OpenAI's guide, changelog and forum post; send 60 messages we wrote in 5 runs; compare right answers, time, cost and order effects
Diagram by Not an AI App.

On 8 October 2026 we read OpenAI's Decisions guide, its developer changelog and its forum announcement. The screenshots are our own captures of those pages, cropped only.

We ran the test the same day with our own OpenAI API key: 60 messages, 5 runs, 300 requests in total, plus one example request. All runs used the default settings, except where we say "thinking off". The whole test cost less than two US cents, paid by the owner of this site. We wrote the 60 messages and chose the right answers ourselves. The charts were made by us. None of the images was made by AI.

This was a small test with short, clear messages. With your own data, the results may be different.