Results
Progress you can inspect.
A result belongs beside its workload, baseline and method. The newest studies come first. Each number in studies 01 to 04 comes from a committed run file.
01
Open models, cold start
Open models on the cold-start task.
We ran three open Hugging Face models on one Mac, on the same 92 cases. An open embedding that reads each pattern's label and description got more cases right than Jev's top option. Research, not a product feature.
Read the study ->92.4%
of 92 cases right (all-MiniLM-L6-v2)
Cross-validated threshold. bge-small-en-v1.5 90.2%, Qwen3-4B 80.4%. Our own run, 8 October 2026.
02
Outcome judge
A small decision model agrees with the reference judge.
On 231 tau2-bench runs, Jev agreed with a strong reference judge more often than Haiku 4.5 did, and it found far more of the failures. Research, not a product feature.
Read the study ->93.9%
agreement with the reference judge
Jev, kappa 0.819. Kev-9B 89.6%, Haiku 4.5 83.5%. Recalculated from the study's row files.
03
Task reward
A transcript judge cannot see a wrong end state.
Against the tau2-bench task reward, four judges were near chance. Each one called most failed runs a success, because a failed run usually reads well. A negative result.
Read the study ->42–49%
balanced accuracy, four judges
240 runs, 120 failed. Each judge called 92 to 106 of the failed runs a success.
04
Cold-start patterns
Find the pattern from the first message.
Two decision models matched a new conversation to its pattern, or to none, far better than an embedding match. They fail in opposite ways. Research, not a product feature.
Read the study ->87.0%
of 92 cases right (Kev-9B)
Jev 85.9%. Embedding at a cosine of 0.5: 55.4%. Recalculated from the study's row files.
05
CRM actions
Better CRM actions, from a better task contract.
A targeted prompt adapter improved Qwen's mean score on seven write-heavy CRM tasks. The improvement came from reading failures and specifying the missing behavior.
Read the study ->+13.1%
relative lift in mean score
0.630 ± 0.087 versus Sonnet 4.6 at 0.557 ± 0.066. Seven tasks; optimized n=10, baseline n=3.
06
Operations workflows
A smaller model. A shorter path to the action.
Output control removed unnecessary generation before sparse fine-tuning repaired remaining failures. A separate Fireworks serving validation measured the resulting 8B route.
Read the study ->5.2×
lower median latency
369 ms versus 1,935 ms. Mean action-level score: 0.9630 versus 1.0000. 90 trajectories per serving comparison.
07
Warehouse labeling
Label the whole table. Inspect the difficult rows.
A post-trained 30B open route labeled 39,962 non-empty comments alongside Sonnet and Opus. The cost comparison is strong; the quality evidence needs a label-by-label reading.
Read the study ->$2.82
reported cost for 39,962 rows
$2.82 open route · $12.48 Sonnet · $139.63 Opus. Agreement between models is not ground-truth accuracy.
