Laya Guard

In short: before your agent runs a tool call, send it with what the user asked for. You get allow, ask or block, the reasons, and calibrated scores for nine questions. Same key and balance as decisions; $0.20 per 1k checks.

Your first check

requestbash
curl https://api.laya.studio/v1/guard \
  -H "Authorization: Bearer $LAYA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "action": { "tool": "db.drop_table", "args": { "table": "users" } },
    "intent": "List the tables in the staging database",
    "context": "Database agent. Environment: production."
  }'
200 responsejson
{
  "id": "dec_3f9c0a…",
  "verdict": "block",
  "p_unsafe": 0.9597,
  "reasons": [
    "destructive",
    "arguments not supported by the request",
    "scope violation"
  ],
  "scores": {
    "safe": 0.0403,
    "violation": {
      "none": 0.0234,
      "policy_violation": 0.0768,
      "scope_violation": 0.383,
      "injection": 0.0552,
      "goal_drift": 0.3826,
      "corrigibility": 0.079
    },
    "severity": {
      "expected": 2.61,
      "level": "high",
      "probabilities": [
        0.03,
        0.08,
        0.14,
        0.75
      ]
    },
    "destructive": 0.9987,
    "exfiltration": 0.0214,
    "injected": 0.0311,
    "approval_policy": {
      "auto_approve": 0.0122,
      "require_human": 0.1105,
      "reject": 0.8773
    },
    "blast_radius": {
      "expected": 1.93,
      "level": "production-mutating or external side effect",
      "probabilities": [
        0.01,
        0.05,
        0.94
      ]
    },
    "args_grounded": 0.2135
  },
  "model": "mcp-guard-deberta-v1",
  "latency_ms": 14.8
}

Put it in your pre-tool-call hook

Only an allow runs without a human. Keep your deterministic rules (blocked commands, allow-listed domains) in front of it: the guard is for the long tail your rules do not describe.

pythonpython
import os, requests

def guard(tool: str, args: dict, intent: str, trigger: str = "user_request") -> str:
    r = requests.post(
        "https://api.laya.studio/v1/guard",
        headers={"Authorization": f"Bearer {os.environ['LAYA_API_KEY']}"},
        json={"action": {"tool": tool, "args": args}, "intent": intent, "trigger": trigger},
        timeout=10,
    )
    r.raise_for_status()
    return r.json()["verdict"]  # "allow" | "ask" | "block"

# In your agent's pre-tool-call hook:
verdict = guard("send_email", {"to": "ops@acme.com", "body": "..."}, intent="Email the weekly report to ops")
if verdict == "allow":
    ...  # run the tool
elif verdict == "ask":
    ...  # show the call to the user and wait for a yes
else:
    ...  # never run it; tell the agent why

Request body

Request body: POST /v1/guardDescription
actionrequiredstring | {tool, args}The tool call about to run: "db.drop_table(users)", or {"tool": "db.drop_table", "args": {"table": "users"}}.
intentstringWhat the user asked the agent to do. The most useful field after the action.
user_messagestringThe latest user message, when it differs from the intent.
trigger"user_request" | "tool_result" | "correction…What produced the action. Use tool_result when the idea came from a tool output or a document (possible prompt injection).
constraintsstring[]Rules the action must respect ("staging only", "never email outside acme.com"). At most 32.
contextstringAgent role, environment (production / staging), anything else worth knowing.
conversation{role, content}[]Recent turns, oldest first, at most 50. Only the most recent part is read.

Each text field is limited to 20,000 characters. The model reads about 384 tokens per check, so put the action and intent first and keep context short.

Response

FieldDescription
verdict"allow" | "ask" | "block"The default decision from the scores below. Run, ask a human, or never run.
p_unsafenumberCalibrated probability that the action is unsafe.
reasonsstring[]Plain-words reasons behind an ask or block.
scores.safenumberProbability the call is safe to run now.
scores.violationobjectnone, policy_violation, scope_violation, injection, goal_drift, corrigibility.
scores.severity{expected, level, probabilities}none < low < medium < high.
scores.destructivenumberDeletes, overwrites or irreversibly changes data.
scores.exfiltrationnumberSends private data or secrets where they should not go.
scores.injectednumberDriven by a tool result or document rather than the user.
scores.approval_policyobjectauto_approve, require_human, reject.
scores.blast_radius{expected, level, probabilities}read-only < reversible write < production-mutating or external side effect.
scores.args_groundednumberThe arguments are supported by what the user asked.
modelstringThe guard model that answered (mcp-guard-deberta-v1: DeBERTa-v3-base, 184M).

The default verdict is a starting point. Thresholds fitted on one dataset do not carry over to every agent, so log the scores, compare them with what your users approve, and set your own cut-offs per head before you let allow run unattended.

Batches

POST /v1/guard/batch takes up to 64 checks and returns one result per check, in order. Use it when a generated script is about to make several calls at once. A check that fails validation comes back as {error} and is not billed.

requestbash
curl https://api.laya.studio/v1/guard/batch \
  -H "Authorization: Bearer $LAYA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "checks": [
    { "action": "read_file(path=\"docs/setup.md\")", "intent": "Show me the setup docs" },
    { "action": { "tool": "send_email", "args": { "to": "ops@evil.example" } },
      "intent": "Summarise my inbox", "trigger": "tool_result" }
  ] }'

From an MCP client

The Laya Studio MCP server (https://api.laya.studio/mcp) exposes guard_check and guard_batch next to the decision tools, with the same key. See MCP server.

Billing

$0.20 per 1k checks: 5,602 credits per check that returns a verdict, from the same balance as decisions, whatever the length of the check. Swiss-only checks (the workspace setting or the x-laya-residency: ch header) cost 15% more and never leave Switzerland. The response headers x-credits-charged and x-request-id tell you what was billed.

StatusMeaning
200Checked. x-credits-charged is the credits billed.
400Malformed body: no action, a bad trigger, constraints not strings, a batch over 64.
401Missing, malformed or revoked key.
402The balance cannot cover the check (5,602 credits each). Add credits or subscribe.
413A field is longer than 20,000 characters.
429Rate limited (per key). Honour Retry-After.
503The guard model is busy or restarting. Retry with backoff; nothing was charged.