Jev vs LLM: Decide or Generate
Many LLM calls are decisions dressed up as text.
When a typed decision beats generated text
Chapters
- Intro
- Decisions dressed as text
- System One vs System Two
- Head to head
- The hybrid pipeline
- When to use which
- Trade-offs and risks
- Recap and checklist
Who it's for
Backend, platform and AI engineers who run LLM calls inside production pipelines (support routing, moderation, approvals, agent tool selection) and are paying generation latency and cost, plus parsing code, for what is really a closed-set classification.
Context
Many LLM calls in production pipelines don't produce text anyone reads. They route, label or approve: the answer is one of a few known options, but it's requested as text and parsed back into a label.
Jev, from TypeSafe AI (out of stealth on 15 September 2026, $40M seed), is a "System One" model in Kahneman's sense: fast judgement instead of slow, written reasoning ("System Two", an LLM). It does not generate text. A request carries a block of state plus one or more typed questions; all questions are evaluated against the state in one parallel pass and each returns a typed answer:
- Choice: one option from your list (up to 255 options), with a
confidence. - Score: a level on your scale (2 to 10 levels), with a
confidence. - Noul: a yes/no question answered as a probability (the probability is the uncertainty signal).
TypeSafe trains it with "Reinforcement Learning for Calibrated Decisions" (RLCD), aiming for confidences that are calibrated: a 90% answer should be right about 90% of the time. That is a vendor claim; independent calibration benchmarks have not been published.
The running example is the Short's support pipeline: every ticket ("Where's my refund?") must go to bill, bug or faq.
Architecture
| Component | Role | Notes |
|---|---|---|
| Ticket | Input | Untrusted natural language |
| Jev (decision model) | Choice(bill, bug, faq, other) + confidence | Answer is always in the set; "other" is the escape hatch |
| Confidence gate | Plain code: if pick != "other" and confidence >= 0.85 | 85% is Correlation One's example threshold; illustrative, tune per question |
| Deterministic workflows | Handle confident, routine cases | No LLM call, no parsing |
| LLM agent | Low-confidence or "other" cases | Reasoning, tools, drafting replies |
| Human tier | Behind the agent | Still-uncertain or high-impact cases |
| Metrics | Log every pick + confidence | Watch the "other" rate, accuracy by confidence bucket, drift |
Request flow
- The ticket text (state) and a Choice question with the allowed queues go to Jev.
- Jev returns one option plus a confidence (tens to hundreds of milliseconds, reported).
- Code logs the pick and confidence, then branches: confident and not "other" → the queue's workflow runs; otherwise → the LLM agent, which can escalate to a human.
The failure modes
1. The LLM classifier (why teams look for an alternative). Code prompts an LLM "Which queue: bill, bug or faq?". The model answers in tokens ("Sure! This sounds like billing."), so the team writes a parser: regex, json.loads, retries. One day the model answers "refunds": sensible, and not in the list. The parser throws, or a default branch silently mis-routes it. Even when parsing succeeds there is no trustworthy confidence to decide when to escalate: token log-probabilities are available on some APIs but are not reliably calibrated after instruction tuning/RLHF, and self-reported confidence is just more generated text. _Fairness note:_ constrained decoding / structured outputs with an enum schema can force an LLM to answer inside the set. That fixes the parsing problem but not the per-call generation latency and cost, or the calibration problem.
2. The decision model is still a model reading untrusted text. Check Point tested Jev on an investment verdict based on a due-diligence report where the attacker controlled one section. Instead of instruction-style injection, they planted plausible evidence (a clean audit opinion from a major firm, a regulatory file number, a revised risk table). The model accepted it as real and the verdict flipped: 59% of attempts succeeded across configurations, at about $0.50 per successful break. Marking the section as untrusted made no meaningful difference. Typed output is an engineering control, not a security boundary.
Why it happens
An LLM is a text generator, so using it as a classifier bolts a parser and a guessed confidence onto a generation call. A decision model removes that mismatch, but it still reads natural language, so anything in its input can steer the decision.
The fix
- Right model for the job: use a decision model as the router, gate or classifier wherever the answer set is closed (routing, triage, moderation labels, screening, approvals, agent tool selection). Keep the LLM for writing, multi-step reasoning, code generation and answer spaces you can't enumerate.
- Cascade on confidence: threshold the confidence in plain code; escalate low-confidence and "other" answers to an LLM agent, and put a human tier behind it (TypeSafe's docs describe a similar tiered pattern: act automatically only at high confidence, gather more data in the middle band, hand off below).
- Design the answer set: include an "other" option and monitor its rate; a missing category otherwise forces confident wrong picks.
- Calibrate on your own data: label a few hundred real items, run both the decision model and the current LLM classifier, and choose the threshold from the accuracy/automation trade-off, per question.
- Treat input as untrusted: screen inputs before they reach the model and keep guardrails and approvals on high-impact actions (Check Point's recommendations).
- Observe: log pick + confidence, sample decisions for review, watch drift, cost and latency for both models.
Trade-offs
- The answer set must be designed and maintained up front; new categories are a schema change.
- A higher threshold means fewer wrong automated picks but more traffic on the slower, costlier LLM/human path.
- Two models to operate, monitor, secure and pay for; the escalation path must scale with the uncertain share.
- The decision model can't explain its answer; if you need a rationale, that's an LLM call.
- Public accuracy and calibration benchmarks are limited, so vendor speed/cost multiples may not hold on your data.
Numbers worth knowing
- Vendor-reported pricing: $0.042 per million input tokens, no output-token charge (TypeSafe; The Register).
- Latency (reported): TypeSafe quotes 70–500 ms end to end; The Register reports ~150 ms typical and ~620 ms in a complex example.
- Vendor-claimed multiples, unverified: more than 100x lower cost and about 200x faster than frontier LLMs on classification (via Correlation One); TypeSafe's own workflow evals claim 193.6x faster and 444.6x cheaper, which TypeSafe itself says are likely on the higher end of real-world gains.
- Security (Check Point): 59% attack success with planted evidence, about $0.50 per successful break.
- Illustrative only: the 85% threshold, the 0.91 confidence in the sequence diagram, and the whole threshold chart (automation rate and wrong-pick rate by threshold).
The router in code
QUEUES = ["bill", "bug", "faq", "other"]
THRESHOLD = 0.85 # illustrative: tune on labelled data
def route(ticket):
ans = decide.choice( # thin wrapper over Jev
state=ticket.text,
question="Which queue handles this ticket?",
options=QUEUES,
)
metrics.observe(ans.pick, ans.confidence)
if ans.pick == "other" or ans.confidence < THRESHOLD:
return llm_agent.handle(ticket) # reason + tools
return WORKFLOWS[ans.pick].run(ticket) # no LLM call
Three question types
Choice
One option from your list, up to 255 options
Score
A level on your scale, 2 to 10 levels
Noul
Yes or no, returned as a probability
Calibrated confidence
Trained so 90% means right 90% of the time (claimed)
Recap: before and after
LLM as classifier
- Output: prose or JSON you parse
- In-set only with enum structured output
- Confidence: logprobs or self-reported
- Seconds per call, pays for output tokens
- Can explain, reason and write
Jev decision model
- Output: a typed value
- Always inside your answer set
- Calibrated confidence (vendor claim)
- ~150 ms, no output charge (reported)
- Cannot write or explain anything
Closed answer set? A typed decision beats parsed prose.
Sources
TypeSafe AI, "Introducing System One Models & Jev" (typesafe.ai blog) — question types, RLCD, latency and pricing claims, workflow evals and their caveat.
Jev documentation as summarised by Flavio Copes, "A deep dive into Jev" — Noul/Choice/Score, 255 options, 2–10 score levels, confidence field, tiered thresholds.
The Register, "TypeSafe AI debuts model for machines that plays Doom" (16 Sep 2026) and "Shut up and calculate: Jev's new AI primitives for coders" (23 Sep 2026) — pricing, ~150 ms / ~620 ms latency, limited public accuracy data.
Correlation One, "What Is an AI Decision Model? Jev, System One Models, and When to Use One Instead of an LLM" — >100x cheaper / ~200x faster claims, 85% threshold example, lack of independent calibration benchmarks.
Check Point Research, "A Decision Model Breaks Like Any Other Language Model: A First Look at Jev" — planted-evidence injection, 59% success, ~$0.50 per break, mitigations.
OpenAI, GPT-4 Technical Report — post-training (RLHF) reduced the calibration of the model's answer probabilities.
OpenAI Structured Outputs / JSON-schema constrained decoding — enum-constrained LLM outputs.
Daniel Kahneman, _Thinking, Fast and Slow_ — System One / System Two.
Coming next in the series: Model Cascades