On this page (10 sections)
Short answer: if you need a calibrated probability per question plus a data-residency story you can put in a contract, Laya Studio is the pick today. If you want the decision model that arrives inside the OpenAI account you already pay for, the Decisions API is the lower-friction bet. As of October 2026, OpenAI has published far less about it than I would want before betting a production pipeline on it. Laya returns a calibrated probability per typed question in one forward pass — roughly 120 ms from Europe — at $5 per million questions after the first 100,000 free.
Key takeaways
- OpenAI entered the decision-only category at DevDay 2026 with the Decisions API, putting it in the same bracket as TypeSafe's Jev and the open-source Laya model rather than alongside chat completions (Flowtivity's three-way comparison).
- Laya Studio generates no text, so it cannot hallucinate a label. You get a number, not a sentence you have to parse.
- We publish our losses: on Banking77 we score 0.425 against Jev's 0.870. That gap is real and it matters for some workloads.
- Laya runs in Switzerland with zero data retention and is Apache-2.0 at the model layer, which is the part compliance teams actually ask about.
- OpenAI's Decisions API pricing, latency and calibration numbers are not in the sources I can cite here. Treat any figure you see elsewhere as unverified until OpenAI publishes it.
What is OpenAI's Decisions API and how does it work?
OpenAI entered the decision-only model category at DevDay 2026 with the Decisions API, which positions it against TypeSafe's Jev and the open-source Laya model rather than against chat completions. Flowtivity's analysis compares the three across cost, latency, calibration and deployment, and describes a three-layer split forming in the market (the Flowtivity breakdown).
The important conceptual move is the output type. A language model returns tokens; a decision model returns a judgement. That is the whole distinction, and it is why the category exists at all — the broader debate about decision models versus language models has been running publicly, including in this Medium piece on the new breed of decision models.
What I cannot tell you is how OpenAI's implementation behaves in production. Their published pricing, latency, and calibration figures are not in the sources I am able to cite. That is not a criticism, it is a fact about the evidence available in October 2026, and it should shape how much of your roadmap you commit this quarter.
How does Laya Studio's System 1 architecture compare?

Laya Studio is a hosted decision API for Laya, an Apache-2.0 System 1 model. You send text or JSON plus typed questions — choice, score, or yes/no — and get a calibrated probability for every question in a single forward pass. No decoding step runs, so there is no generated label to hallucinate.
"System 1" here means a model that produces a judgement in one pass instead of reasoning in generated text. I wrote about why that matters for classification pipelines in the future of AI classification beyond LLMs.
Two architectural consequences worth naming. First, because the model emits probabilities rather than a string, you can set thresholds per question and change them without touching the model. Second, because the weights are open, you can inspect what you are depending on. Neither property is unique to us, but they are the ones teams tell me they were missing when they were regex-parsing LLM output.
The honest limitation: Laya does not accept image input. Text and JSON only. If your pipeline needs to classify a scanned form or a photo, this is not the tool.
Performance: latency and calibration benchmarks in 2026

Laya returns a calibrated probability per question in one forward pass, roughly 120 ms measured from Europe. Calibrated probability means the number reflects how often the model is right — a 0.8 should be correct about 80% of the time, not just "confident". That property is what makes a threshold defensible in an audit.
Where we lose: on the Banking77 intent benchmark we score 0.425 against TypeSafe Jev's 0.870. I would rather you read that from us than discover it after migration — the numbers and the setup are in our Jev comparison. Jev's genuine advantage is exactly this: on fine-grained intent labelling it is the stronger classifier, and if Banking77-style granularity is your core task, Jev should be on your shortlist.
If latency is the constraint rather than accuracy, the regional picture matters more than the model — I covered that in optimizing decision APIs for low-latency European applications. OpenAI's latency figures for the Decisions API are not published in the sources available to me, so I will not put a number next to ours.
What does the cost of an AI classification API look like?
Laya Studio prices per question: 100,000 free, then $5 per million. That is arithmetic you can budget against — a million classifications a month costs about $5, ten million costs about $50, and the cost scales linearly with volume rather than with token length. The detail is on Laya Studio's pricing page.
The reason per-question pricing matters more than it sounds: with LLM-prompt classification, cost tracks input length. Long customer emails cost more than short ones. Per-question pricing removes that variable, which makes forecasting a support or moderation budget meaningfully easier.
OpenAI's Decisions API pricing is not published in the sources I can cite, so I am not going to guess at it or draw a comparison line that would be fiction. Check their current published rates directly before you model anything.
Comparison table: OpenAI Decisions API vs Laya Studio
Everything below is either our own published specification or sourced to the Flowtivity analysis. Where a cell says "not published", I could not find a citable figure as of October 2026 — that is a gap in the evidence, not a claim about capability.
| Dimension | OpenAI Decisions API | Laya Studio |
|---|---|---|
| Category | Decision-only model, launched at DevDay 2026 (Flowtivity) | Hosted decision API for the Apache-2.0 Laya model |
| Output | Not published in cited sources | Calibrated probability per typed question; no generated text |
| Question types | Not published | choice, score, yes/no |
| Latency | Not published | ~120 ms, one forward pass, from Europe |
| Pricing | Not published | 100,000 free, then $5 per million questions |
| Data residency | Not published | Switzerland, zero data retention |
| Open weights | Not published | Yes, Apache-2.0 |
| Protocol | Not published | Compatible with TypeSafe Jev's /v1/systemone |
| Languages | Not published | English plus 100+ |
| Image input | Not published | Not supported — text or JSON only |
What about GPT-6 Luna classification, image input, and context window?
I have no citable figures for GPT-6 Luna classification, so I am not going to invent them. If you are evaluating it, ask the vendor for calibration data and a reproducible benchmark setup rather than a headline accuracy number — that is the only kind of claim worth acting on.
On the dimensions I can answer: Laya supports English plus 100+ languages, takes text or JSON input, and answers typed questions about that input. It does not take images. We do not publish a token-window figure alongside this article, so treat any number you see attributed to us as unverified unless it is on laya.studio.
The practical point is that "context window" is a language-model framing. A decision model does not need to hold a conversation; it needs to answer specific questions about a payload you hand it. Ask vendors for the input size they support, not the window size they advertise.
Use cases where each decision API excels
Laya fits pipelines where the output has to be a defensible number: support ticket routing, escalation triage, moderation flags, and fraud signals where you tune a threshold rather than accept a label. The pattern for routing is written up in real-time AI classification for customer support routing.
The Decisions API's genuine advantage is distribution. If your stack already runs on OpenAI, it arrives inside an account, SDK surface and billing relationship you already have — one fewer vendor, one fewer contract, one fewer thing to explain to procurement. That is a real benefit and I would not dismiss it.
Where I would pick Laya over it: regulated European workloads, anything where a wrong label has to be explained to an auditor, and high-volume classification where per-question pricing keeps the bill flat.
Migrating from LLMs to a specialized decision API
The migration is usually smaller than teams expect, because the hard part is not the model call — it is the parsing layer you built around it. If you are currently prompting an LLM and regexing the answer into a label, you already have the question schema. You just stop parsing.
If you are coming from TypeSafe Jev, Laya is protocol-compatible with /v1/systemone, so in many cases it is a base-URL change rather than a rewrite. Our Jev alternatives roundup covers the switching details.
Three things to check before you cut over:
- Thresholds. Decide the probability cutoff per question before launch, and decide who owns changing it.
- Calibration on your data. Run your own labelled sample. Our Banking77 number is ours, not yours.
- Residency and retention. Confirm in writing what is stored and where, then keep that answer for the audit.
Choosing the right decision model: a practical framework
Work through this in order and the choice usually makes itself.
- Do you need a probability or a label? If a threshold decision, you need calibration. If you just need text back, you may not need a decision API at all.
- Does the decision have to be explainable? Regulated or audited decisions favour a model that returns a number with a known calibration curve.
- Where must the data live? If the answer is "inside Switzerland" or "nowhere after the response", that narrows the field fast.
- What is your volume? Per-question pricing rewards volume; per-token pricing punishes long inputs.
- How fine-grained is the label set? If it looks like Banking77, test Jev first. We score 0.425 there and we say so.
If you want to test the first four against real traffic, you can start on the free tier at laya.studio — 100,000 questions, no card.
FAQ
Is OpenAI's Decisions API cheaper than Laya Studio?
I cannot answer that from the sources available to me — OpenAI's Decisions API pricing is not published in the material I can cite as of October 2026. Laya is $5 per million questions after 100,000 free. Check OpenAI's current published rates before modelling.
What is a calibrated probability in AI?
It is a probability that reflects how often the model is actually right. If the model says 0.9, roughly nine out of ten of those predictions should be correct. Uncalibrated confidence scores look precise but do not support a defensible threshold.
Does Laya Studio support image input?
No. Laya accepts text or JSON plus typed questions. If your classification pipeline needs to read images or scanned documents, you need a different component in front of it.
Can I switch from TypeSafe Jev to Laya Studio without rewriting my client?
Often, yes. Laya is compatible with Jev's /v1/systemone protocol, so for many integrations the change is the base URL and credentials. Verify your specific request shapes against the docs before cutover.
Which decision API is better for a regulated European workload?
Laya Studio, on the residency dimension: it runs in Switzerland with zero data retention and the model weights are Apache-2.0. That said, residency is only one axis — if your task needs fine-grained intent accuracy, benchmark both on your own labelled data first.
Topics
- OpenAI Decisions API vs Laya Studio
- System 1 model comparison
- OpenAI GPT-6 Luna classification
- Laya decision model features
- real-time AI classification API
- AI decision API latency
- cost of AI classification API
- calibrated probabilities in AI
