Guide

The Developer's Guide to Integrating System 1 AI for Real-time Decision Making

A code-first guide to integrating System 1 AI decision APIs: latency, cost, typed questions, calibrated probabilities, and production monitoring.

Laya Studio9 min read
On this page (9 sections)

Sending one piece of text to a System 1 decision API returns a calibrated probability for every typed question you attach, in a single forward pass — roughly 120 ms from Europe, at $5 per million questions after the first 100,000 free. To integrate System 1 AI, you send text or JSON plus typed questions to a classification endpoint and read back probabilities. There is no prompt to engineer, no generated label to hallucinate, and no output string to parse. Most teams have a working call within an hour. The hard part is deciding what to do with the number.

Key takeaways

  • A System 1 decision API classifies; it does not generate. That single design choice removes the parsing layer most LLM classification setups need.
  • Pricing is per question, not per request — so your cost lever is asking fewer, better questions, not batching requests.
  • Calibrated probabilities are only useful if you validate thresholds on your own data first. Laya scores 0.425 on Banking77 against TypeSafe Jev's 0.870, and that gap is real.
  • Zero data retention means your audit trail lives on your side. Log inputs and outputs yourself.
  • Keep the protocol layer thin. Laya speaks TypeSafe Jev's /v1/systemone shape, which makes switching a config change rather than a rewrite.

What Is a System 1 Decision API and Why Should You Integrate It?

A System 1 decision API is a hosted endpoint that returns a calibrated probability for each question you attach to a piece of text, in one forward pass, without generating any text. "System 1" borrows the psychology term for fast, automatic judgement: the model classifies rather than reasoning out loud, so there is no free-text label to hallucinate and no string to parse.

That matters because most teams currently classify by prompting a generative model and then writing regex to clean up the answer. You get a label, a mood, and a support ticket when the model replies "Sure! Based on the text, this is probably Billing." A decision API returns a number instead, and numbers compose.

The integration itself is ordinary API work: authenticate, send a payload, handle a response, retry on failure. If you have ever wired up a payments or email provider, you already know the shape. IBM's overview of API integration covers the general patterns — connection, transformation, error handling — and MuleSoft's API integration resources go deeper on the architectural side if you are designing this across several services.

How Do You Choose the Right System 1 API for Real-time Decisions?

How Do You Choose the Right System 1 API for Real-time Decisions?

Judge candidates on five things: output type, latency, calibration evidence, data residency, and pricing model. Everything else is secondary. A fast API that returns an uncalibrated score you cannot threshold is slower in practice than a slower one you can act on directly.

CriterionQuestion to askWhy it bites you later
Output typeProbability per question, or free text?Free text forces a parsing layer you must maintain
LatencyRound-trip from your region, under loadPublished latency is usually measured from somewhere convenient
CalibrationIs there a benchmark, and what is the setup?Uncalibrated scores make thresholds arbitrary
Data residencyWhere does inference run, and is data retained?Retrofitting residency after launch is expensive
PricingPer request, per token, or per question?Token pricing is unpredictable at volume
ProtocolIs the request shape portable?Lock-in is a schema, not a contract

For low-latency AI integration, measure from your own infrastructure rather than trusting a number on a marketing page. The ~120 ms figure for Laya is measured from Europe; from other regions you should expect more, and you should test it. There is a longer discussion of this in optimizing AI decision APIs for low-latency European applications.

How Do You Integrate Laya Studio With Your Application?

How Do You Integrate Laya Studio With Your Application?

The flow is three steps: build a payload with your text and typed questions, POST it, read probabilities keyed by question id. Below is the shape of a call — check the current API docs for exact field names before you ship, since schemas move faster than blog posts.

Python

python
import os, requests

API = os.environ["LAYA_API_URL"] # from your dashboard
KEY = os.environ["LAYA_API_KEY"]

payload = {
 "text": "My invoice was charged twice and nobody replied for a week.",
 "questions": [
 {"id": "intent", "type": "choice",
 "options": ["billing", "shipping", "account", "other"]},
 {"id": "needs_human", "type": "yes_no"},
 {"id": "urgency", "type": "score", "min": 0, "max": 1},
 ],
}

r = requests.post(API, json=payload,
 headers={"Authorization": f"Bearer {KEY}"},
 timeout=2.0)
r.raise_for_status()
probs = r.json()["probabilities"]
# {"intent": {"billing": 0.91, ...}, "needs_human": 0.77, "urgency": 0.84}

Node.js

js
const r = await fetch(process.env.LAYA_API_URL, {
 method: "POST",
 headers: {
 "Authorization": `Bearer ${process.env.LAYA_API_KEY}`,
 "Content-Type": "application/json",
 },
 body: JSON.stringify(payload),
 signal: AbortSignal.timeout(2000),
});
if (!r.ok) throw new Error(`decision API ${r.status}`);
const { probabilities } = await r.json();

Two things worth noticing. The timeout is short on purpose — a decision API that has not answered in two seconds is not helping your request path. And the questions travel with the request, so the same text can answer several decisions in one round trip.

How Should You Structure Data Input and Typed Questions?

Ask short, mutually exclusive questions, and pick the question type that matches the decision you are actually making. Choice questions give you a probability per option; yes/no gives you one number; score gives you a bounded value. Asking three overlapping questions about the same concept produces three correlated numbers and no extra information.

  • Keep option labels disjoint. "Billing" and "payment issue" will fight each other and split probability mass.
  • One decision per question. If you need routing and escalation, that is two questions, and you pay for two.
  • Trim boilerplate. Email signatures, legal footers, and quoted reply chains add noise.
  • Test in your languages. Laya supports English plus 100+ languages, but a threshold tuned on English is not automatically right for German.
  • Version your question set. Changing an option label silently changes your output distribution.

How Do You Interpret Calibrated Probabilities in Your Application Logic?

A calibrated probability is one where the number reflects how often the model is right — if you act on everything scoring above 0.8, roughly 80% of those should be correct. That is the whole promise, and it is what makes a threshold defensible to an auditor rather than a magic constant someone picked in a meeting.

In practice, use three bands rather than one cutoff:

  1. Auto-act above your high threshold (routing, tagging, auto-reply).
  2. Abstain in the middle — send to a human queue, or ask a clarifying question.
  3. Auto-act the other way below your low threshold, where that is meaningful.

Validate the thresholds on your own labelled sample before launch. Do not assume a model that wins a public benchmark wins yours: on Banking77, Laya scores 0.425 against TypeSafe Jev's 0.870. That is a real loss and it is worth knowing before you build on it. If your domain looks like Banking77, benchmark both. There are worked examples of threshold design in 5 real-world applications of calibrated probabilities in AI decision making.

What Are the Best Practices for Production Deployment and Monitoring?

Treat the decision API like any other hard dependency: timeout it, retry it, monitor it, and have a fallback. The difference from a database is that failures are silent — a bad probability still looks like a number, so nothing throws and nothing pages you.

  • Set an explicit timeout (1–2 seconds) and a circuit breaker. Fall back to a human queue rather than a default label.
  • Retry only idempotent calls, and never retry on a 4xx.
  • Log inputs, questions, probabilities, and the action taken on your side. Zero retention means Laya keeps nothing, so your audit trail is your responsibility.
  • Run in shadow mode first. Compute decisions, log them, act on none of them, compare against human labels for a week.
  • Watch the distribution, not just the error rate. A drift in the share of high-confidence decisions is your earliest signal.
  • Count questions, not calls. At $5 per million questions with 100,000 free, cost is driven by how many questions you ask per text.

What Are the Most Common Integration Problems and How Do You Fix Them?

Most integration failures are not model failures. They are timeouts, auth mistakes, and question design that produces a useless spread of probabilities.

SymptomLikely causeFix
Every probability near 0.5Ambiguous or overlapping questionsRewrite options so they are disjoint
Timeouts under loadConnection setup per request, no poolingReuse a session/client, raise the timeout slightly
401 / 403Key in the wrong header or wrong environmentCheck the auth header against docs, rotate keys per environment
Good scores, bad outcomesThreshold never validatedLabel a sample and measure precision above the cutoff
Cost higher than expectedToo many questions per textCut to the two or three that drive a decision

How Do You Future-Proof Your Real-time Decision Architecture?

Keep the model behind a thin internal interface, version your question sets, and keep the protocol portable. If your application calls classify(text, questions) and that function happens to POST to a Jev-compatible /v1/systemone endpoint, swapping providers is a config change rather than a refactor.

Two pressures are worth planning for now. Data residency requirements keep tightening, and running inference in Switzerland with zero retention is one way to answer them — see the Swiss advantage for AI data residency and zero retention. And the EU AI Act's August 2026 deadline has already passed for high-risk systems, so if your classification feeds a regulated decision, your logging and threshold documentation are now compliance artifacts, not nice-to-haves. That is covered in navigating the EU AI Act's August 2026 deadline.

If you are weighing this against TypeSafe Jev specifically, Laya Studio vs TypeSafe Jev lays out the benchmark differences honestly, including the ones Laya loses.

If you want to test the integration path described here, you can start free at laya.studio — 100,000 questions, no card, and a call you can have running before lunch.

FAQ

What is a System 1 decision API in plain English?

It is a hosted endpoint that reads your text and returns a probability for each question you ask about it — for example, "is this billing?" — without writing any sentences back. You get numbers instead of prose, which means you can threshold them, log them, and act on them automatically.

How long does it take to integrate a classification API?

For a single service, usually an afternoon: credentials, one POST call, and a threshold. The longer work is deciding which questions drive a real decision and validating your thresholds against labelled examples from your own traffic.

Do I need an SDK to use a decision API?

No. A plain HTTP client is enough, and it is often better — fewer dependencies, and you can see exactly what is on the wire. An SDK helps mainly when you want typed request and response objects in a large codebase.

What happens to my data when I send it to a decision API?

With Laya Studio, nothing is retained: inference runs in Switzerland with zero data retention. That also means you cannot query the provider for past decisions, so build your own logging if you need an audit trail.

Can I switch from TypeSafe Jev without rewriting my code?

Usually yes, if you kept the protocol layer thin. Laya is compatible with Jev's /v1/systemone request shape, so the change is typically credentials and endpoint rather than a new integration.

Topics

  • integrate System 1 AI
  • System 1 AI implementation
  • real-time decision API integration
  • AI classification API setup
  • developer workflow AI
  • low-latency AI integration
  • System 1 model deployment
  • API best practices AI

Live demo

Reading is good. Trying is better.

See real answers on five example messages, with a calibrated probability for every answer in about a tenth of a second. Sign up and your first 5 runs on your own messages are free.

Example answer, captured live

Answered in Switzerland

“Hi, I was charged twice for order #4821 ($129.00). Please refund the duplicate charge before Friday, our books close then. This is the second billing mistake this quarter and we're starting to look at other vendors.”

What does the customer want?

  • refund100%
  • other<0.1%
  • cancel0%
Is it urgent: 17.7%Might they leave: 30.3%

0 words generated · 3 questions in one pass · 354 ms round trip when captured