Stop Generating, Start Deciding: The Bounded-Decision Pattern OpenAI Just Productised, And Why Most Of Your Agent Loop Should Use It
The quiet announcement at DevDay was a Decisions API: a constrained version of GPT-6 Luna that takes context, a question and a finite set of allowed answers and returns exactly one, in well under a second, for classification, routing and choosing an agent's next step. It is in limited preview and the docs are thin, but the idea behind it is the most useful thing a senior engineer can take from the week - because it names a pattern that separates agent systems that work in production from ones that demo well. Most steps in an agent loop are not creative. They are decisions from a small set: which tool, which route, continue or stop, safe or unsafe, confident or escalate. Treating those as bounded decisions - enumerated answers, constrained output, calibrated confidence, deterministic consequences - makes a loop faster, cheaper, testable and auditable. This is the education piece: what the pattern is, when generation is still right, how to implement it with any model today, and how it reshapes an agent architecture for finance.
AlchmAI Engineering14 min read
1 of N
What a bounded decision returns: exactly one answer from a set the developer defined, never free text
<1 sec
Latency class of OpenAI's previewed Decisions API, roughly an order of magnitude faster than asking the same model to write
~80%
Of steps in a typical production agent loop that are decisions, not generation - tool choice, routing, stop/continue, safe/unsafe, escalate
0
Parsing failures when the output is a constrained enum rather than a paragraph the code has to interpret
Every agent framework shows the same loop: observe, think, act, repeat. What the diagrams hide is that most of the 'think' steps are not thinking in any creative sense. The agent has to pick which tool to call next from a list of ten. It has to decide whether the task is done. It has to classify an input as safe or suspicious, a request as in-scope or out, a result as confident or needing a human. Those are decisions from small, known sets, and in most production systems they are made by asking a model to write a paragraph and then parsing it. That is slow, expensive, brittle and impossible to measure properly.
OpenAI's Decisions API, previewed at DevDay on 29 September, productises the alternative. You give it context - text or images - a question, and the allowed answers. A constrained GPT-6 Luna returns one of them, fast. The pseudocode in early write-ups is telling: context is a support ticket and account tier, the question is 'which approved workflow owns this case', the answers are a list of workflow names, and the calling code checks the answer is in its own allowed set before acting. The API is in limited preview with thin documentation and unpublished pricing. The pattern it names is available today with any capable model, and it is worth adopting now.
Why It Matters More Than It Sounds
- 01Speed. A decision is a handful of output tokens. The same model that takes two seconds to write a plan returns a label in a fraction of that, and a small model does it faster still. Loops with ten decisions a task feel different when each is 100 milliseconds.
- 02Cost. Output tokens are the expensive ones. A loop that generates reasoning at every step can cost ten times one that decides at most steps and generates at two.
- 03Testability. A bounded decision has a right answer. You can build a labelled set, measure accuracy and calibration per decision, and gate releases on them - exactly as you would a classifier, because that is what it is.
- 04Auditability. 'The model chose route B with confidence 0.93 under policy v12' is a log line a compliance team can read. A paragraph of model reasoning is not.
- 05Safety. The model cannot answer outside the set. It cannot invent a tool, a desk, a route or an action. The whole class of 'the agent did something nobody defined' goes away for those steps.
Implementing It Today, With Any Model
You do not need to wait for a preview. A strict tool schema whose fields are enums gives you the constraint; a short system prompt gives you the framing; and a calibration pass gives the confidence meaning. The example below is the next-step decision for a reconciliation agent.
export const NEXT_STEP = ["fetch_ledger", "fetch_settlements", "compare", "draft_exception_report", "escalate_to_human", "done"] as const;
export const RISK = ["none", "low", "high"] as const;
export const DECIDE_NEXT = {
name: "decide_next_step",
description: "Choose the single next step for the reconciliation agent. Choose ONLY from the enum.",
input_schema: {
type: "object",
properties: {
next_step: { type: "string", enum: [...NEXT_STEP] },
risk: { type: "string", enum: [...RISK], description: "high if the step would touch money, client data or external systems" },
confidence: { type: "number", minimum: 0, maximum: 1 },
reason_code: { type: "string", enum: ["missing_data", "data_ready", "mismatch_found", "all_matched", "ambiguous", "policy"] },
},
required: ["next_step", "risk", "confidence", "reason_code"],
},
strict: true,
} as const;
export async function decideNext(state: AgentState, llm: LLM, threshold: number) {
const r = await llm.call({
system: "You are the controller of a reconciliation agent. Decide the next step using only the tool.",
messages: [{ role: "user", content: JSON.stringify(state.summary()) }], // a compact state, not the transcript
tools: [DECIDE_NEXT], maxTokens: 60,
});
const d = r.toolInput as { next_step: string; risk: string; confidence: number; reason_code: string };
if (d.confidence < threshold || d.risk === "high") return { ...d, next_step: "escalate_to_human" }; // deterministic consequence
return d;
}from collections import Counter
def evaluate(decisions, labelled):
"""decisions: {case_id: (answer, confidence)}; labelled: {case_id: correct_answer}.
Reports accuracy, per-class confusion and calibration so each decision point gets a release gate."""
correct = sum(1 for k, (a, _) in decisions.items() if labelled.get(k) == a)
confusion = Counter((labelled[k], a) for k, (a, _) in decisions.items() if labelled.get(k) != a)
buckets = {}
for k, (a, c) in decisions.items():
b = min(int(c * 10), 9)
n, ok = buckets.get(b, (0, 0))
buckets[b] = (n + 1, ok + int(labelled.get(k) == a))
calibration = {f"{b/10:.1f}": (ok / n if n else None) for b, (n, ok) in sorted(buckets.items())}
return {"accuracy": correct / max(1, len(decisions)), "top_confusions": confusion.most_common(5), "calibration": calibration}
# Gate: accuracy >= 0.97 on the labelled set, no confusion between any step and 'done',
# and calibration monotonic above the escalation threshold.When Generation Is Still The Right Tool
- Explaining a decision to a human: the label says what, the paragraph says why. Generate the why only when a person will read it.
- Drafting: reports, summaries, messages, code. The output is language and the reviewer is a person.
- Open-ended extraction where the schema cannot be known in advance - though even here, a typed schema with optional fields beats free text.
- Genuine planning in novel situations. Keep it rare, keep it logged, and follow it with bounded decisions that check the plan against policy.
What It Does To An Agent Architecture
Adopting the pattern changes the shape of the system. The controller becomes a sequence of bounded decisions with deterministic consequences, each with its own labelled set and release gate. Tools are chosen by enum, not by the model inventing a call. Safety checks - is this input an injection attempt, is this action in scope, does this need approval - become fast classifiers run before every consequential step rather than hopes expressed in a system prompt. And generation is pushed to the edges: the first-step understanding of a messy input, and the last-step explanation to a human. In a bank, that is also the architecture a model-risk team can validate, because each decision point looks like the classifiers they already know how to assess.
def run(task, state, decide, tools, generate, threshold):
"""A controller where every branch is a bounded decision and generation happens at the edges."""
state.intent = generate.understand(task) # generation: messy input -> typed intent
for _ in range(state.max_steps):
if decide.is_injection(state.latest_observation()): # bounded: {clean, suspicious}
return state.halt("suspicious_input")
d = decide.next_step(state, threshold) # bounded: enum + calibrated confidence
if d.next_step == "escalate_to_human":
return state.escalate(d.reason_code)
if d.next_step == "done":
return generate.explain(state) # generation: for the human reader
if decide.needs_approval(d.next_step, state): # bounded: {auto, approve}
state.await_approval(d.next_step)
continue
state.observe(tools[d.next_step](state)) # deterministic consequence
return state.escalate("max_steps")“The agents that survive production are not the ones that reason the most. They are the ones that decide the most and reason only where a human will read the result.”
The Bottom Line
OpenAI's Decisions API preview gives a name and a product to a pattern senior engineers should already be applying: most steps in an agent loop are bounded decisions from enumerated sets, and they should be implemented as constrained classifications with calibrated confidence and deterministic consequences, not as generated text to be parsed. The result is faster, an order of magnitude cheaper, testable against labelled sets, auditable as log lines and structurally safer because the model cannot answer outside the set. Generation stays where language is the product: understanding a messy input and explaining an outcome to a person. You can build it today with strict tool schemas on any capable model and swap in a dedicated decisions endpoint when one suits. That is how we architect agents for financial firms as an AI agency in London, and it is the most practical idea to come out of DevDay.
References & Further Reading
- Hugging Face blog - What is OpenAI Decisions API? A practical guide. huggingface.co/blog/sora-2/what-is-openai-decisions-api-a-practical-guide
- InfoQ - OpenAI DevDay 2026 recap for developers. infoq.com/news/2026/10/openai-devday-2026
- Learnetto - OpenAI DevDay 2026: every announcement. learnetto.com/openai-devday-2026-announcements
- Anthropic Docs - Tool use with Claude (strict schemas). docs.anthropic.com/en/docs/build-with-claude/tool-use
- OpenAI Docs - Structured outputs. platform.openai.com/docs/guides/structured-outputs
- Guo et al. - On Calibration of Modern Neural Networks. arxiv.org/abs/1706.04599
- Simon Willison - OpenAI DevDay 2026 live blog. simonwillison.net/2026/Sep/29/openai-devday-2026-live-blog
AlchmAI Engineering
Engineering, London
Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.
Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information