On 2026-09-15 TypeSafe AI opened early access to Jev, which it calls a System One model 1. It reads text like an LLM and does not generate any. You send a state (the trace) and typed questions; it returns probabilities over answers you defined 2. For an eval team that is an LLM judge with the critique removed and the price cut by two orders of magnitude. Figures are as of October 2026, for jev-1.13.0, the version both studies below tested.
What a decision model is
An LLM judge writes its verdict as text, usually after a critique, and you parse it. A decision model never generates. Jev answers three question types: Noul (a 0 to 1 probability that a statement is true), Choice (one option from your list), and Score (a level on an ordered rubric) 2, the last two with a probability per option or level 3. Choice and Score also return a confidence from 0 to 1, computed from the shape of that distribution: 1 when all the mass sits on one answer, 0 when it is spread evenly 3. It measures spread, not correctness; the vendor says calibration is measured across groups of predictions and does not guarantee that any single answer is correct 2.
Two consequences follow. The output is schema-valid by construction, so parse failures and invented labels disappear; the launch post's "can't hallucinate" means that and no more 1. And there is no rationale. System One models do not explain their reasoning 2, and DeepEval's JevEval metric builds its reason field in code, with no LLM call 4. A verdict you cannot read is one your open coding loop cannot learn from.
The cost and latency arithmetic
Vendor-reported, as of October 2026: $0.042 per million input tokens, output free, and rate limits of 40 requests and 100,000 tokens per second that TypeSafe says can change without notice 5. The launch post quotes 70 to 500 ms end to end and concedes its 444.6x cheaper, 193.6x faster headline figures are "on the higher end of real world gains" 1.
Independent measurements come in under the vendor's headline but the gap stays large. Li et al. measured 12.182 and 1.89 s for GPT-6 6. Rao and Callison-Burch found three flash-tier LLM judges cost 16 to 325 times as much and took 28 to 350 times as long 7. On the distilled judges assumptions (1,200 input tokens per judgment, 100,000 traces a day), list price is 5 a day. The binding limit is 40 requests per second: at one judgment per request, about 3.5 million a day.
What changes is coverage. A 1 to 5 percent sample finds a regression in a 0.5 percent segment slowly. At $5 a day you score every trace and keep sampling for the layers that still cost real money: the frontier judge and the human labeler.
What the independent evidence says
Li et al. tested Jev against sixteen judges with blinded human adjudication 6; Rao and Callison-Burch against three flash-tier LLMs on nine human-labeled rubric panels 7. Accuracy, in percent:
| Workload | Jev | Comparator | Study |
|---|
| Ordinary preference (RewardBench) | 92.5 | 92.5 (GPT-6) | Li |
| Grounded factuality (HaluEval) | 87.3 | 88.4 (GPT-6) | Li |
| Hard correctness (JudgeBench) | 78.6 | 93.1 (GPT-6) | Li |
| Logic puzzles, pooled | 68.4 | 95.9 (GPT-6) | Li |
| Wrong answer written more elaborately (RM-Bench hard pairs) | 76.6 | 90.1 (GPT-6) | Li |
| Reference-free prose responses | 53.5 | 56.0 (GPT-5.4) | Li |
| Binary checklist criteria (RiceChem) | 81.0 | 76.1 to 79.2 (three flash LLMs) | Rao |
| Graded essay scale, exact level (ELLIPSE) | 13.4 (15.3 as a Score) | 29.8 (Gemini 3.8 Flash) | Rao |
Where the verdict can be read off the text, Jev is within about three points of GPT-6 6, and on binary checklist criteria it is ahead of flash-tier judges or level with them 7. Where the judge must derive or check a result (expert knowledge, code, math, logic), it trails GPT-6 by 7.0 to 27.6 points 6. On graded scales it led on only two of seven panels, both times asked as a Score 7. On JudgeBench it was 6.4 points more accurate when the preferred answer came second, so run pairwise inputs in both orders 6. The derive-or-check gap is the one every judge has.
Confident errors break naive cascades
The obvious design is a cascade: accept Jev's verdict above a confidence threshold, escalate the rest to a frontier judge. With a 0.9 threshold frozen in advance, Li et al.'s cascade (each pair judged in both orders, probabilities averaged) escalated 31.5 percent of 1,610 held-out pairs, scored 93.4 against GPT-6's 92.5, and cost 41.4 percent of GPT-6's fee 6. On JudgeBench it escalated about 65 percent, so the savings shrink where the judge is weakest. Their threshold is on the probability of the chosen label, not on TypeSafe's confidence field, so a 0.9 in the paper is not a 0.9 in your code; tune your own cut on whichever number you gate on.
The problem is what the gate lets through. On Jev's 84 most confident errors across the seven graded-scale panels, 242 of 252 LLM verdicts (96.0 percent) repeated the same wrong answer, against 50.3 percent expected under independence, and even with oracle thresholds no cascade beat the best single judge by more than 2.7 points 7. Li et al. saw it from the other side: on the 2,332 items where Jev put all its probability on one label, Jev and GPT-6 were right or wrong together on all but 2 6. Escalation fixes the errors Jev is unsure about; on the confident ones a second LLM mostly agrees with it. Only a check that errs differently catches those: an executable verifier or a human.
The gate is an injection target
Jev is sold for guardrails too 1, and LangChain's TypeSafe integration includes an experimental middleware that asks it whether a tool call is risky and refuses the call if so 8. A classifier that decides whether a tool runs is something attackers will write text for. TypeSafe is candid: Jev does not treat its input as hostile, and an injected instruction or text that argues for its own classification can move the answer 9. Its RAG cookbook says its own injection filter is not a security boundary 10. Three rules follow.
- Never use one model as both gate and judge. If Jev blocks tool calls in production and also grades those traces, a payload that flips the gate flips the eval, and the dashboard shows a clean pass on the exact traffic that got through. Grade the gate with a different model family and a human-labeled attack set from the prompt injection taxonomy.
- Keep untrusted text out of the gate's input. TypeSafe tells you to send only the fields a question needs 9. For a tool-call gate, send the user's request, the proposed call, and the tool description, not raw tool output or retrieved documents, where indirect injection arrives. The LangChain docs warn that tool arguments and conversation state go to TypeSafe 8, so check what that state holds before you trust the gate.
- Keep a human on irreversible actions. The LangChain middleware refuses risky calls but never asks for approval; its docs say to pair it with human-in-the-loop middleware 8.
The architecture this site recommends
flowchart TD
A[Every trace] --> B[Decision model<br/>typed verdict + confidence]
B -->|at or above threshold| C[Accepted verdict]
B -->|below threshold| D[Frontier LLM judge<br/>verdict + critique]
A -->|1-5% random slice| D
C -->|audit sample| E[Human gold labels]
D --> E
E --> F[TPR and TNR per slice<br/>accuracy by confidence bin]
Decision model on every trace. Frontier LLM judge on everything below the threshold plus a random slice, which is where your critiques for error analysis come from. Human labels on a sample of both piles, since the confident errors live in the accepted one. Score both judges against one gold set as in calibrating your judge against humans: TPR and TNR per slice, not agreement percent.
Re-verify calibration on your own distribution, and again when the model changes. The jev-latest alias moves with each release; TypeSafe says to pin a versioned ID once you have tuned thresholds 5. DeepEval warns that a metric scored in its system_one mode is a different number from the same metric under llm, so keep each metric on one mode across runs 11.
What to do this week
- Rewrite your judge criteria as Noul or Choice questions, run them on your gold set (100 labeled traces minimum) at a pinned version, and report TPR and TNR per slice beside your frontier judge.
- Bin gold-set accuracy by confidence and pick the threshold where accuracy above it meets your bar. Slices where it escalates more than half the traces stay with the frontier judge or an executable check.
- Add a weekly human audit of accepted high-confidence verdicts. Escalation will never surface those errors.
- If any classifier gates tool calls, give its eval a different judge, strip tool output and retrieved text from its input, and point your injection suite at the gate itself.