Operant uses LLM judges in two places. The eval gate uses a judge to compare a cheaper model with your baseline. Each conversation also gets an outcome score. The A/B split shows that score for each variant. LLM judges are slow, and each verdict costs money.
So we asked a research question: can a small decision model do the task of the outcome judge? This post gives the results. This work is research. It is not part of the product.

The models
We compared four judges on the same runs, with the same question and the same options:
- Jev is a hosted decision model from TypeSafe, version
jev-1.13.0. It writes no text. It answers a typed question with a probability for each option. - Kev-9B is an open variant of Jev on Hugging Face. We ran it on our own Mac with MLX. No text left the network.
- Haiku 4.5 is a small LLM judge.
- Sonnet 5 is a stronger LLM judge. We used it as the reference.
The question was "Did the agent do what the user asked?" The options were success, failure and unclear.
The runs
The material is 240 runs from tau2-bench, an open benchmark of agents in airline and retail tasks. Half of the runs passed and half failed by the benchmark reward. Each judge read each full transcript.
Result 1: Jev agrees with the reference judge
Sonnet 5 gave a verdict on 231 of the 240 runs. We used its verdict as the label:
- Agreement: Jev 93.9%, Kev-9B 89.6%, Haiku 4.5 83.5%.
- Balanced accuracy: Jev 89.8%, Kev-9B 79.3%, Haiku 4.5 66.3%.
- Recall on failures: Jev 82.4%, Kev-9B 60.8%, Haiku 4.5 35.3%.
- Cohen's kappa: Jev 0.819, Kev-9B 0.683, Haiku 4.5 0.515.
Jev found 82.4% of the failures that the reference found. Haiku found 35.3%.
Jev also gives a probability. So you can let Jev decide only when it is sure, and send the other runs to an LLM judge. At a floor of 0.7, Jev decided 85.7% of the runs alone. It agreed with the reference on 98.5% of those.
Result 2: no judge sees the benchmark reward
The benchmark reward checks the final state of the task. It checks the database and the policy, not only the words. Against that reward, every judge was near chance:
- Balanced accuracy: Jev 49.2%, Kev-9B 42.1%, Haiku 4.5 45.8%, Sonnet 5 47.5%.
- Recall on failed runs: Jev 20.0%, Kev-9B 10.8%, Haiku 4.5 8.3%, Sonnet 5 21.7%.
- AUROC: Jev 0.536, Kev-9B 0.465, Haiku 4.5 0.531, Sonnet 5 0.572.
Each judge called most of the 120 failed runs a success. A failed run usually reads well. The user says thanks, but the agent wrote the wrong change to the database or broke a policy. A judge that reads only the transcript cannot see that error.
What this means for outcome scores
An outcome score from a transcript tells you if the user goal looks met. It does not prove that the final state is correct. This matters when a system picks successful conversations to teach a skill.

The next research step is to run a decision model in shadow beside the LLM judge. That step is not shipped.
The numbers
Each number in this post comes from the committed row files of the study. We calculated them again from those rows. The full tables are on the results page: outcome judge against the reference and judges against the benchmark reward.
