OpenAI Decisions API for Product Managers: Faster Routing at 10x Speed
TL;DR
OpenAI announced the Decisions API at DevDay 2026: a specialized endpoint that uses the Luna model to answer bounded, finite questions from text or image context. It is currently 10x faster than Luna via the standard chat API and is positioned for classification, request routing, and selecting an agent's next action. If pricing comes in low enough, it could reshape every AI pipeline that currently calls a full LLM just to make a routing decision.
The AI PM Minute
One tactic to make you a sharper AI PM, twice a week. 60 seconds to read. Free.
No fluff. Unsubscribe anytime.
What the Decisions API Actually Does
The Decisions API, currently in limited preview, is a purpose-built endpoint for one type of problem: picking from a predefined set of answers. You give it context (text, an image, or both), a question, and a fixed list of valid answers. It returns one of those answers. No open-ended generation, no chain of thought, no streaming tokens.
The model underneath is Luna, OpenAI's fast, cost-optimized tier from the GPT-6 family. But calling it through the Decisions API is reportedly 10x faster than calling Luna through the standard chat completions endpoint. OpenAI attributes this to the constrained output format: when the model only has to select from a fixed vocabulary of answers, the inference path is substantially shorter.
Input
Text context, image context, or both. The question you are asking. A defined list of valid answers (the constrained output space).
Output
One answer from the list you provided. No explanation, no generation, no tokens beyond the answer itself.
Speed
10x faster than Luna via the standard API in OpenAI's benchmarks. Exact latency numbers have not been published but sub-100ms inference is implied for short contexts.
Use cases OpenAI named
Content classification, request routing, and choosing an agent's next action in a multi-step workflow.
Current availability
Limited preview as of DevDay 2026. No pricing announced publicly. You can join the waitlist through the OpenAI developer portal.
The product positioning is clear: this is not a replacement for the chat completions API. It is an inference shortcut for the subset of AI problems that are really just fast, reliable classification.
Why 10x Faster Matters for AI Product Design
A 10x latency reduction changes the architectural calculus in two specific situations.
First: anywhere you currently call a full LLM just to make a routing decision. Plenty of production AI pipelines do exactly this. A user submits a query; you call GPT-4o mini to classify whether the query is a question, a command, a complaint, or nonsense; then you route accordingly. The classification call adds 300 to 800ms of latency and costs tokens you are not using for anything generative. If the Decisions API handles that classification in under 100ms, the routing step becomes close to invisible to the user.
Second: inside agentic pipelines where the agent makes many small decisions per task. A browser agent completing a 20-step workflow might need to decide at each step which tool to call next, whether the current page matches the expected state, or whether to abort and request human input. If each of those micro-decisions costs 600ms today, the full workflow takes 12 seconds longer than it needs to. Decisions API could compress that significantly.
Where latency is already acceptable
Batch pipelines, overnight processing, backend enrichment jobs. For these, raw speed is not the constraint. Cost per call matters more, and pricing has not been announced.
Where latency directly affects UX
Real-time chat routing, live moderation, step selection in interactive agents, low-latency voice AI. This is where 10x speed compounds into a noticeably better experience.
Where open-ended generation is required
Decisions API cannot help. If the answer is not in a predefined list, you need chat completions or a reasoning model. Do not try to force generation tasks into a constrained output format.
Where reliability beats speed
High-stakes classification (medical triage, legal document routing, financial fraud flagging). Here you need explainability and audit trails that a single constrained output does not provide.
Four Use Cases Worth Building On
These four patterns map well to what the Decisions API is designed for. Each one currently works with a regular LLM call, but would benefit from the speed and cost reduction the Decisions API provides.
1. Intent Router
Problem it solves: A user sends a message. Before the main LLM processes it, you need to classify the intent: question, task request, complaint, off-topic. Each intent routes to a different prompt or agent.
Decisions API fit: Classic classification with a small fixed label set. Give the Decisions API the message and a list of 4 to 8 intent categories. Sub-100ms routing means the rest of the pipeline starts faster.
Watch out for: Decisions API cannot explain its answer. If you need to log why a message was routed a certain way for debugging or compliance, you will need to add a separate step.
2. Agent Next-Action Selector
Problem it solves: A multi-step agent needs to decide at each step which tool to call next, given the current state and history. Today this is often a full LLM call with a large system prompt listing available tools.
Decisions API fit: Each decision is bounded: the agent can only take one of N predefined actions. Pass the current state as context and the action list as the answer set. Compresses per-step latency across long workflows.
Watch out for: Works best when the action space is stable. If your tool list changes frequently, you will need to update the valid answers on each call, which adds implementation overhead.
3. Content Moderation First Pass
Problem it solves: User-generated content must be screened before display or processing. A full LLM call adds latency and cost to every submission.
Decisions API fit: Classify each piece of content as safe, review-needed, or block. The Decisions API handles this as a three-way classification. Speed matters here: moderation that delays posting by 300ms frustrates users at scale.
Watch out for: For content with legal implications (CSAM, threats, financial fraud), a bounded classification should feed into a human review queue, not act as the final gate alone.
4. Form and Document Triage
Problem it solves: Incoming forms, support tickets, or documents need to be triaged to the right team or workflow before a human or downstream agent processes them.
Decisions API fit: The Decisions API accepts image context, which means it can classify documents by their visual type (invoice, contract, ID, screenshot). Combined with text context, it handles mixed-media triage efficiently.
Watch out for: Document classification with high accuracy needs a well-defined answer set. Vague categories like 'other' or 'unknown' will accumulate and create downstream problems if the Decisions API does not surface its confidence.
Build on Emerging APIs With Confidence
The AI PM Masterclass covers how to evaluate new APIs and make architectural decisions before a technology is fully documented. Taught live by a former Apple and Salesforce Sr. Director PM.
How Pricing Will Determine Whether This Changes Architecture
OpenAI has not announced pricing as of DevDay 2026. That single variable will decide whether the Decisions API is a niche speed optimization or a fundamental shift in how AI pipelines are designed.
The argument for low pricing is strong. Luna is already OpenAI's cost tier; the Decisions API requires even less compute per call because the output is constrained. If OpenAI prices it below Luna per call, and especially if it prices per-decision rather than per-token, the economics of adding a routing step to every AI pipeline become very different.
Pricing above Luna per call
Niche use case. Teams only adopt it where latency is already a critical constraint and they cannot achieve the same result with a smaller prompt on a standard endpoint.
Pricing at parity with Luna per call
Moderate adoption. Teams gradually replace classification LLM calls with Decisions API calls when rebuilding pipelines. No forcing function to refactor working pipelines.
Pricing below Luna, billed per decision not per token
Architectural shift. Every pipeline that currently uses an LLM for routing rethinks that decision. Routing layers become cheap enough to run on every request without budgeting concern.
A secondary pricing question is the image context surcharge. Multimodal inputs cost more on every endpoint today. If the Decisions API charges significantly more for image classification than for text classification, the document triage use case becomes less attractive.
What This Means for Your AI Product Roadmap
The Decisions API does not require you to rearchitect anything today. It is in limited preview, pricing is unknown, and reliability at scale has not been demonstrated publicly. The right posture for most product teams is monitor and prepare rather than migrate immediately.
Audit your current LLM calls for classification patterns
Find every place in your pipeline where you are calling an LLM and the output is really just a label or a selection from a fixed set. These are the candidates. Document them now so you can evaluate the Decisions API against each one when pricing is announced.
Do not redesign pipelines until pricing is live
The architectural benefits are real, but they only materialize at the right price point. Refactoring a working pipeline for a speed optimization that costs more per call than what you replaced is not a trade worth making.
Watch the accuracy benchmarks carefully
Luna via the standard API has known accuracy characteristics on classification tasks. The Decisions API adds the constrained output format, but that does not guarantee higher accuracy. Evaluate on your own domain before committing to it in production.
Plan for limited auditability
A constrained output gives you no reasoning trace. For regulated industries or high-stakes routing decisions, you will need a secondary logging strategy: store the input, the output, and a timestamp. The Decisions API will not explain itself.
One structural implication is worth internalizing now: the Decisions API signals that OpenAI sees a category of AI work that should not require a full generative model. That framing, whatever its pricing, is likely to persist and be adopted by other providers. The pattern of "fast bounded decisions separate from open-ended generation" is probably where the ecosystem is heading.
Make Smarter Bets on Emerging AI APIs
The AI PM Masterclass teaches you how to evaluate new APIs and make architectural decisions before documentation is complete. Stop waiting for the ecosystem to settle before you act.
Related Articles
Before you go: get the AI PM Minute
One tactic to make you a sharper AI PM, twice a week. 60 seconds to read. Free.
No fluff. Unsubscribe anytime.