Results / Outcome judge
A small decision model agrees with the reference judge.
We asked if a small decision model can do the task of an LLM outcome judge. On 231 agent runs, Jev agreed with a strong reference judge 93.9% of the time. Haiku 4.5 agreed 83.5% of the time.
Study run 28 September 2026 · Numbers recalculated from the row files on 8 October 2026 · Research, not a product feature
Jev (jev-1.13.0)
93.9%
agreement · kappa 0.819
Kev-9B (open variant)
89.6%
agreement · kappa 0.683
Haiku 4.5
83.5%
agreement · kappa 0.515
Agreement with the reference verdict on 231 runs · an unclear verdict counts as a disagreement
The workload and the label.
The material is 240 runs from tau2-bench: airline and retail tasks, with agents on claude-3-7-sonnet and gpt-4.1. By the benchmark reward, 120 runs passed and 120 failed. Each judge read each full transcript and answered one question: did the agent do what the user asked? The options were success, failure and unclear.
The reference is Sonnet 5 with the outcome prompt of Strata. It gave a verdict on 231 of the 240 runs: 180 success and 51 failure. On this page its verdict is the label. The reference is a stronger LLM, not ground truth.
| Metric | Jev (jev-1.13.0) | Kev-9B (open variant) | Haiku 4.5 |
|---|---|---|---|
| Agreement | 93.9% | 89.6% | 83.5% |
| Balanced accuracy | 89.8% | 79.3% | 66.3% |
| Recall on reference-success | 97.2% | 97.8% | 97.2% |
| Recall on reference-failure | 82.4% | 60.8% | 35.3% |
| AUROC of P(success) | 0.982 | 0.916 | 0.959 |
| Cohen's kappa | 0.819 | 0.683 | 0.515 |
231 runs where the reference decided (180 success, 51 failure). For an LLM judge, P(success) is its stated confidence in a success verdict, or one minus its confidence in a failure verdict.
A cascade: the small model decides when it is sure.
Jev gives a probability for each option. A cascade lets Jev decide when its top probability is at or above a floor, and sends the other runs to Haiku 4.5. At a floor of 0.7, Jev decided 85.7% of the runs alone and agreed with the reference on 98.5% of those.
| Floor | Jev decides | Agreement on what Jev decides | Cascade agreement (rest to Haiku) |
|---|---|---|---|
| 0.6 | 94.4% | 96.3% | 93.5% |
| 0.7 | 85.7% | 98.5% | 91.8% |
| 0.8 | 74.5% | 98.8% | 90.0% |
| 0.9 | 58.9% | 100.0% | 89.6% |
| 0.95 | 42.4% | 100.0% | 87.0% |
Haiku 4.5 alone agrees with the reference on 83.5% of the runs.
Method.
Jev is a hosted decision model from TypeSafe. It writes no text and returns a probability for each option. Kev-9B is its open variant on Hugging Face (jaredpalmer/kev-9b), run locally with MLX in bf16 on an M3 Ultra. All judges saw the same sample, the same question, the same rendering and the same options.
Every number on this page comes from the committed row files of the study (one row for each run and judge). We recalculated each metric from those rows with an independent script and kept the run log.
Limitations and source record.
- The label is another LLM judge, not a human verdict. Agreement with it is not accuracy against ground truth.
- There are only 51 reference failures, so the failure recall figures have wide uncertainty.
- All runs come from one benchmark in two domains. Other workloads can give other results.
- Each Jev verdict sends the transcript to TypeSafe. Kev-9B runs locally.
- Against the benchmark reward itself, every judge is near chance. See the companion study.
Source: the upstream Strata research study of 28 September 2026 (docs/jev-evaluation, row files at commit ada801f). Recalculated by Operant on 8 October 2026. This is research evidence, not live telemetry and not a product feature.
