Cinematic Short Decision Models 56 sec Why use it

Jev vs Laya

Same idea. One closed, one open.

Which to pick

Few labels, fast, on-prem: Laya. Many labels, zero ops: Jev. Open-ended work: an LLM.

Context

"System One" decision models answer a typed question (a Choice from your list, a Score on your scale, or a yes/no probability) with a confidence, instead of generating prose. They compete for the decision layer of an AI pipeline (routers, gates, classifiers), not with chat LLMs.

  • Jev (TypeSafe AI, launched 15 September 2026): closed, proprietary, a hosted managed API. TypeSafe markets its confidences as calibrated (vendor claim).
  • Laya (Convai Innovations, released 18 September 2026; hosted API at laya.studio): open weights under Apache-2.0. The English checkpoint is a ~421M-parameter ModernBERT-large encoder with a decision head; a ~322M multilingual variant covers 100+ languages. It runs on your own GPU or CPU, and a community demo even ran it in a browser tab with ONNX Runtime Web (about 0.8–1.3 s per decision on CPU there).

The episode reuses the series' ticket-routing example: every support ticket must be routed to bill, bug or faq.

Architecture

ComponentRoleNotes
TicketInput"Where's my refund?"
JevChoice(bill, bug, faq) + confidence via a hosted APIClosed, managed; network round trip on every call
LayaSame typed question, run on your own hardwareOpen weights, Apache-2.0; you operate and can fine-tune it
Typed pick`bill \bug \faq` + confidenceAlways inside the set, so code branches on it directly
Threshold / LLM (not drawn)Escalation for low-confidence or open-ended casesAs in the "Jev: Decisions, Not Text" episode; 85% threshold illustrative

Request flow

  1. The ticket and a typed question ("Which queue: bill, bug or faq?") go to the decision model.
  2. Jev: an HTTPS call to TypeSafe's service. Laya: a local forward pass on your GPU/CPU (or a call to Laya Studio).
  3. The model returns one option from the list plus a confidence.
  4. Code branches on it; low-confidence or open-ended tickets go to an LLM or a human.

The advantages

  1. Shared: decisions, not text. Both return a typed value inside the set you defined; no prose parsing.
  2. Laya: control and cost. Open weights, Apache-2.0, data never leaves your environment, no per-call fee, fine-tunable.
  3. Laya: latency (Laya-claimed). The model card reports ~32.8 ms median for one question on a Tesla T4 GPU; third-party runs measured Jev 1.13.0 at ~236–276 ms p50 end-to-end over the internet. Not the same setup (local GPU vs network API; Laya Studio's hosted API adds a round trip too), so the "~8x" gap is a claim, not a like-for-like benchmark.
  4. Jev: accuracy on large label sets. On Banking77 (77 intents), third-party results show Jev 0.870 vs base Laya 0.425; Laya's default option budget struggles past ~20 options. A Laya checkpoint fine-tuned on a benchmark's training split reached 0.766 on a typed-decisions set, where Jev scored 0.727 zero-shot.
  5. Jev: zero ops. Managed service, long inputs (64k context reported), strong zero-shot baseline.

Where it fits (and where it doesn't)

  • Laya: few labels (under ~20), tight latency budgets, high volume, privacy or on-prem requirements, or when you can fine-tune on labelled data.
  • Jev: many labels, long inputs, no appetite for running models, or a managed contract.
  • Neither: open-ended writing, multi-step reasoning, code. Keep an LLM there.
  • Known weaknesses: Laya's calibration needs temperature fitting on your data (reported ECE 0.466 before, 0.081 after), and one write-up found the English model at 0.080 accuracy on Bengali script while reporting 0.945 confidence. Jev is a closed box with enterprise pricing and a network hop; its out-of-the-box ECE was reported at 0.246 in one comparison.

Trade-offs

  • Laya: you own serving, scaling, monitoring and recalibration; accuracy drops sharply on big label sets unless fine-tuned.
  • Jev: per-call cost (The Register reported $0.042 per million input tokens, no output charge), vendor lock-in, and data leaving your network.
  • Latency numbers come from different setups; measure both on your own traffic before choosing.

Numbers worth knowing

  • Laya ~421M parameters (English), ~322M multilingual; Apache-2.0 (Laya model card / Laya Studio).
  • Laya ~32.8 ms median per question on a T4 GPU (Laya-reported). Jev ~236–276 ms p50 end-to-end (third-party measured; TypeSafe states 70–500 ms). Shown on screen as "Laya claims 33 ms on a GPU. Jev's API: about 250 ms".
  • Banking77: Jev 0.870, base Laya 0.425 (third-party comparison). Shown as 87% vs 43%.
  • Sources: Laya Studio blog "Laya vs. Jev" and compare page; DEV Community "Jev vs Laya: the same AI idea, one closed and one open" (jamilxt); wilsonwu.me "Jev vs Laya"; AlphaMatch "Jev vs Laya"; Vishal Mysore's browser demo on Medium.

Coming next in the series: Model Cascades

Found this useful?

Subscribe for the next episode, or share it with the person who owns this part of your stack.

Keep going