Results / Task reward

A transcript judge cannot see a wrong end state.

We compared four outcome judges with the task reward of tau2-bench on 240 runs. Every judge was near chance. Each one called most failed runs a success, because a failed run usually reads well.

Study run 28 September 2026 · Numbers recalculated from the row files on 8 October 2026 · A negative result

Balanced accuracy

42.1–49.2%

four judges · 50% is chance

AUROC of P(success)

0.465–0.572

0.5 is chance

Failed runs called success

92–106

of 120 failed runs

240 tau2-bench runs · 120 passed and 120 failed by the benchmark reward

The workload.

The material is 240 runs from tau2-bench: airline and retail tasks, with agents on claude-3-7-sonnet and gpt-4.1. The benchmark reward checks the final state of each task, such as the database and the policy rules. By that reward, 120 runs passed and 120 failed.

Each judge read each full transcript and answered one question: did the agent do what the user asked? The options were success, failure and unclear.

MetricJev (jev-1.13.0)Kev-9B (open variant)Haiku 4.5Sonnet 5
Decided (not unclear)100.0%95.8%91.2%96.2%
Accuracy on decided49.2%43.9%50.2%49.4%
Balanced accuracy49.2%42.1%45.8%47.5%
Recall on passed runs78.3%73.3%83.3%73.3%
Recall on failed runs20.0%10.8%8.3%21.7%
AUROC of P(success)0.5360.4650.5310.572
Failed runs called success96 / 120106 / 12098 / 12092 / 120

Balanced accuracy is the mean of the recall on passed runs and the recall on failed runs. For an LLM judge, P(success) is its stated confidence in a success verdict, or one minus its confidence in a failure verdict.

Why every judge fails here.

A failed run usually reads like a success. The user says thanks. But the agent wrote the wrong change to the database, or it broke a policy. The transcript does not show that error, so a judge that reads only the transcript cannot find it.

This limit applies to the stronger model too. It matters when a system picks successful conversations to teach a skill: a transcript score shows that the user goal looks met, not that the final state is correct.

Method.

The judges are Jev (a hosted decision model from TypeSafe), Kev-9B (its open variant on Hugging Face, run locally with MLX), Haiku 4.5 and Sonnet 5. All judges saw the same sample, the same question, the same rendering and the same options.

Every number on this page comes from the committed row files of the study. We recalculated each metric from those rows with an independent script and kept the run log.

Limitations and source record.

  • One benchmark in two domains, with two agent models. Other workloads can give other results.
  • The judges answered a question about the user goal. The reward checks policy and final state. The gap between those two questions is part of the result.
  • Against a strong reference judge, the same judges agree far more. See the companion study.

Source: the upstream Strata research study of 28 September 2026 (docs/jev-evaluation, row files at commit ada801f). Recalculated by Operant on 8 October 2026. This is research evidence, not live telemetry and not a product feature.

Continue the investigation.