OpenAI Decisions API: Verdicts Instead of Essays, Explained

N
Navs
Published on October 11, 20267 min read
Tags:decisions-apiopenaigpt-6-lunaai-agentsapi
OpenAI Decisions API: Verdicts Instead of Essays, Explained

OpenAI spent years teaching models to write better essays. On October 6, 2026, it shipped an API that skips the essay entirely. The Decisions API, opened to all developers in public beta, takes your text or images, asks it a few bounded questions, and returns typed answers your code can branch on — a probability, a choice, or a score. No prose to parse, no JSON schema to babysit, and, per OpenAI, about 10x faster than asking the same question through the Responses API.

It lands in a category that already has a breakout star: TypeSafe AI's Jev, the "System One" decision model that spread through Vercel's AI Gateway in September. Here's what the Decisions API is, what it costs — and five tools already putting decision-style models to work.

What OpenAI shipped

Announced on OpenAI's developer forum on October 6, 2026: the Decisions API, live in public beta at POST /v1/decisions, so apps can "choose the right model, tool, or action in near real time."

Decisions API screenshot

Three things define it. First, the endpoint returns answers, not text: you bundle your input — text, or messages mixing text and images — with a set of named questions, and get back an answers array keyed by those names. Second, one model powers it at launch: gpt-6-luna, the efficiency-tuned member of the GPT-6 family. Third, the price: $0.10 per million input tokens, and you pay only for input — no output-token, cache-read, or cache-write charges. (Regional processing premiums and long-context multipliers can still apply.)

OpenAI says the endpoint produces typed answers about 10x faster than the Responses API, and expects general availability "in the coming weeks." That 10x figure is the company's own measurement — no independent latency study had been published at the time of writing. The beta ships with Zero Data Retention support and HIPAA eligibility for qualifying customers, plus US and EU (EEA/Switzerland) data residency. A limited preview had appeared at DevDay on September 29.

The three answer types

Everything the API does is built from three question types:

TypeWhat you askWhat comes back
Predicate"Is this statement true?"A probability from 0 to 1
Choice"Which option fits best?"The chosen option, a probability for every option, plus confidence
Score"Where does this fall on my rubric?"A probability-weighted position on your ordered levels, with confidence

The guide's worked examples are concrete: a damaged-product photo checked with a predicate, a billing complaint routed to a department with a choice, a bug scored for severity with a score. Multiple questions about the same input are evaluated in parallel in one request.

Two constraints matter. Images must be inline base64 data URLs — hosted links and uploaded file IDs are refused, so URL-passing pipelines need a fetch-and-encode step. And the beta has no input caching: developers on the announcement thread confirmed it, which makes long repeated contexts pricier than on the Responses API until that changes.

What it costs — and the Jev comparison

Input-only billing is the pricing headline. The same gpt-6-luna through the Responses API costs $0.10 per million input tokens plus $0.50 per million output tokens, so a Decisions call drops the output bill entirely — a big deal for classification workloads with short answers.

OpenAI Decisions APITypeSafe Jev 1.13gpt-6-luna on Responses
Input price / 1M tokens$0.10$0.042$0.10
Output priceFreeFree$0.50 / 1M
ImagesYes (inline base64)No — text onlyYes
Answer typesPredicate, choice, scoreNoul (boolean), choice, scoreFree text / your JSON schema
StatusPublic beta (Oct 6, 2026)Early access, via Vercel AI GatewayGenerally available

Jev is the reason this category exists. TypeSafe AI's flagship "System One" model launched out of stealth in mid-September 2026, spread through Vercel's AI Gateway — nearly 13% of paid Gateway teams tried it in the first 24 hours, per Vercel — and announced an $870M raise at a $7.5B valuation on October 9 (Andreessen Horowitz leading, per the company's own blog). Jev is cheaper per token ($0.042 vs $0.10) and, by TypeSafe's figures, answers in 70–500ms — but it's text-only, while OpenAI's version reads images and sits inside an account most teams already have. Neither vendor has published a head-to-head accuracy benchmark; one developer's few-hundred-question test found Jev more accurate on nuanced calls while Luna made about three times as many confident wrong answers — but the author called it "a strong early signal rather than a general benchmark." Treat calibration as unresolved.

Five ways to put decision models to work

The ecosystem moved fast. Five tools already working with decision-style models, from live-data pipelines to agent frameworks.

1. SerpApi: decisions over live search data

SerpApi screenshot

SerpApi's team published one of the first end-to-end Decisions API tutorials: a remote job finder. The pattern is general — fetch live listings with SerpApi's google_jobs engine, send them to the Decisions API with your preferences as questions, and filter out non-matching roles. They had earlier explored the same approach with a Jev-based fact checker over search results. If your decisions need fresh web data, copy this pairing: a search API for the facts, a decision model for the verdict.

2. Strands Agents: route agent work to the right model

Strands Agents screenshot

Strands Agents, AWS's open-source SDK for building AI agents, inspired a community tutorial within days of the beta: plug the Decisions API in as a classifier behind Strands' ModelRouter, so each request is judged against the candidate models' descriptions and routed to the small cheap model or the capable one, with a confidence threshold deciding. Not every request needs your strongest model, and a typed choice is a cleaner router than a paragraph of reasoning.

3. OpenRouter: one endpoint, thirteen decision models

OpenRouter screenshot

OpenRouter's alpha /decisions endpoint exposes the pattern across thirteen models — including gpt-6-luna and typesafe/jev-1.13 — behind one request shape. To benchmark OpenAI's offering against Jev without wiring up two vendors, this is the fastest route: same request, different model parameter, measured side by side.

4. Jagent: browser Jev agents you can use today

Jagent screenshot

Jagent (jev-agent.com) is an independent, unofficial guide to Jev that ships working Jev agents in the browser: paste an essay for six rubric scores with probabilities, check a resume against a job description, or call the same graders over HTTP and MCP inside Claude and Cursor. The site also fronts multiple decision models behind an OpenAI-compatible endpoint, and its benchmark pages publish the independent tests that went against Jev alongside the ones that favored it — unusual honesty.

5. TypeSafe AI's Jev: the benchmark to beat

Jev screenshot

The model that started it all. Jev 1.13 lists $0.042 per million input tokens with free output, answers in a claimed 70–500ms, and handles up to 64k tokens per request — text only, no open weights, closed managed API, early access via waitlist and gateways. If you're evaluating the Decisions API, Jev is the baseline your tests should include: cheaper and faster on TypeSafe's numbers, but blind to images and outside your OpenAI account.

Who should use it

This is a developer API. The natural users are teams burning chat-model calls on classification work: support ticket routing, content moderation queues, lead scoring, relevance checks, RAG reranking, and agent internals — anywhere the pattern is "prompt for a paragraph, then regex out a label."

Strengths and limits

The strengths are real: typed outputs your code can branch on, input-only billing that undercuts chat calls for classification, image support Jev lacks, and enterprise checkboxes (ZDR, HIPAA, EU residency) that suit regulated moderation flows.

The limits deserve equal airtime. It's a beta: one model, no caching, fields or limits may move before GA. The 10x speed claim is OpenAI's own and unverified. Community testing has surfaced calibration quirks — a loaded-coin test found choice questions putting nearly all probability mass on the most likely outcome rather than estimating the true posterior, and rounding a score can name a rubric level the model gave 0% probability. OpenAI's own advice stands: set thresholds from your own labeled examples, weighted by the cost of false positives against false negatives.

How to start

Try a predicate, a choice, and a score in the official Playground (platform.openai.com/decisions), then check your SDK version before calling POST /v1/decisions. Budget from the $0.10-per-million-input-token base rate, and keep your Responses API path warm as a fallback — a preview can change limits or pricing without notice. Choosing between vendors? Run your own labeled set through both the Decisions API and Jev first: the right answer depends on your data, not the vendors' multipliers.

Bottom line

The Decisions API is OpenAI's bet that the next phase of applied AI is not better prose but better verdicts — bounded, typed, machine-readable answers for routing, triage, and moderation. It's cheaper and cleaner than parsing chat output, and the ecosystem is already building on it. But it's a beta with one model, unverified speed claims, and a cheaper rival the community keeps benchmarking it against. Try it on your own labeled data this week.

References

Share this article