Results / Cold-start patterns

Find the pattern from the first message.

A new conversation must find its pattern from the first user message, or say none. Two decision models got 85.9% and 87.0% of 92 cases right. An embedding match at a cosine of 0.5 got 55.4%.

Study run 28 September 2026 · Numbers recalculated from the row files on 8 October 2026 · Research, not a product feature

Kev-9B (open variant)

87.0%

of 92 cases · top option

Jev (jev-1.13.0)

85.9%

of 92 cases · top option

Embedding, cosine ≥ 0.5

55.4%

of 92 cases

A case is right when the matcher picks the expected cluster, or none when no cluster fits

The cases.

The material is 54 conversations from a demo workload. 38 belong to a family that Cortex clustered. For these, the matcher saw four clusters, each with its label and description, and a none option. 16 conversations expect none: 7 are sanctions screening, which is near SAR triage but has no cluster, and 9 are one-off tasks.

The study also asked 38 more questions. For each family member, it removed the cluster of that family, so the right answer was none. That gives 92 cases.

MatcherFamily members rightNone cases rightLeave-one-out: said noneAll cases right
Jev, top option92.1%37.5%100.0%85.9%
Jev, assign only at p ≥ 0.9 (floor set after the results)89.5%87.5%100.0%93.5%
Kev-9B, top option71.1%93.8%100.0%87.0%
Kev-9B, assign only at p ≥ 0.9 (floor set after the results)18.4%100.0%100.0%66.3%
Embedding, cosine ≥ 0.521.1%31.2%100.0%55.4%
Embedding, cosine ≥ 0.56 (tuned on this set)21.1%93.8%100.0%66.3%
Embedding, nearest centroid (no threshold)21.1%0.0%0.0%8.7%

38 family members, 16 none cases and 38 leave-one-out cases. The embedding compares the first message with each cluster centroid (text-embedding-3-small).

Two ways to be wrong.

Jev and Kev-9B have almost the same total, but they fail in opposite ways. Jev found 92.1% of the family members, but it forced near misses into the nearest cluster: it sent all 7 sanctions-screening conversations to SAR triage. Kev-9B found fewer family members, 71.1%, and its errors were mostly none.

For routing, a none is the safer error. It leaves the call on its usual model. A wrong pattern sends the call to a skill for a different task. A probability floor helps Jev, but a real floor must come from shadow traffic, not from this small set.

Method.

Jev is a hosted decision model from TypeSafe. Kev-9B is its open variant on Hugging Face (jaredpalmer/kev-9b), run locally with MLX. Each decision model answered one choice question over the cluster options and none, and gave a probability for each option.

Every number on this page comes from the committed row files of the study. We recalculated each metric from those rows with an independent script and kept the run log.

Limitations and source record.

  • 54 conversations from one demo workload. The cases are few, and the families are synthetic.
  • The 0.9 floors and the 0.56 embedding threshold were chosen after the results, so they are optimistic.
  • The embedding compares a raw first message with centroids built from goal summaries. A different embedding input can give a different result.

Source: the upstream Strata research study of 28 September 2026 (docs/jev-evaluation, row files at commit ada801f). Recalculated by Operant on 8 October 2026. This is research evidence, not live telemetry and not a product feature.

Continue the investigation.