Results / Open models, cold start

Open models on the cold-start task.

We ran three open models from Hugging Face on one Mac, on the same 92 cases as the first cold-start study. An open embedding model that compares the first message with the label and description of each pattern got 92.4% right. An open 4B instruct model got 80.4%.

Run on 8 October 2026 on an Apple M4 Max · Every number comes from the run files · Research, not a product feature

all-MiniLM-L6-v2

92.4%

of 92 cases · cross-validated threshold

bge-small-en-v1.5

90.2%

of 92 cases · cross-validated threshold

Qwen3-4B-Instruct-2507

80.4%

of 92 cases · top option

A case is right when the matcher picks the expected cluster, or none when no cluster fits

The task.

The cases are the 92 cases of the first cold-start study: 38 family members, 16 conversations that expect none, and 38 leave-one-out cases. We rebuilt each first message from the same demo recordings, and checked that the ids, the passes and the expected answers equal the rows of that study.

The options are the four Cortex clusters, each as its label and description, and "None of these tasks." The question is the same: "Which of these known tasks is the user asking the assistant to do?"

MatcherFamily members rightNone cases rightLeave-one-out: said noneAll cases rightSanctions sent to SAR
Qwen3-4B-Instruct-2507, top option86.8%25.0%97.4%80.4%7 of 7
Qwen3-4B-Instruct-2507, assign only at p ≥ 0.986.8%25.0%97.4%80.4%7 of 7
all-MiniLM-L6-v2, nearest option (no threshold)97.4%0.0%0.0%40.2%6 of 7
all-MiniLM-L6-v2, cosine ≥ 0.3298 (tuned on this set)94.7%81.2%97.4%93.5%0 of 7
all-MiniLM-L6-v2, cross-validated threshold92.1%81.2%97.4%92.4%0 of 7
bge-small-en-v1.5, nearest option (no threshold)94.7%0.0%0.0%39.1%6 of 7
bge-small-en-v1.5, cosine ≥ 0.5982 (tuned on this set)89.5%87.5%97.4%92.4%0 of 7
bge-small-en-v1.5, cross-validated threshold86.8%87.5%94.7%90.2%0 of 7
First study: Jev, top option92.1%37.5%100.0%85.9%7 of 7
First study: Kev-9B, top option71.1%93.8%100.0%87.0%0 of 7
First study: text-embedding-3-small to the centroids, cosine ≥ 0.521.1%31.2%100.0%55.4%—

A cross-validated threshold is tuned on the other 53 conversations only, for each conversation. The first-study rows come from its recalculated row files.

What we found.

The text that the embedding reads matters. The first study compared the first message with the centroid of the goal summaries in each cluster, and found 21.1% of the family members. Here, two small open embedding models compared the message with the label and description, the same text that the decision models read. With a cross-validated threshold, they found 92.1% and 86.8%. We changed two things at once, the model and the text, so this run does not show which change matters more.

A threshold is what lets an embedding say none. With no threshold, the nearest option is never none, so the none cases score 0.0%. With a threshold, the two embedding models got 81.2% and 87.5% of the none cases right, and sent no sanctions-screening conversation to SAR triage.

The 4B model was sure of itself, and sometimes wrong. Qwen3-4B gave its top option a probability of 0.9 or more on 91 of the 92 cases, so the 0.9 floor changed nothing. Like Jev, it sent all 7 sanctions-screening conversations to SAR triage.

Method.

Qwen3-4B-Instruct-2507 ran in 4-bit MLX format, with one forward pass for each case. The options were letters. The probability of an option is the probability of its letter as the next token, normalised over the letters. The embedding models ran with sentence-transformers and compared each first message with each "label: description" by cosine. bge-small-en-v1.5 got its query instruction.

Every number on this page comes from the run files: one row for each case and model, the metrics, and a run log with the command, the model revisions and the timestamps.

Limitations and source record.

  • Cortex wrote the cluster descriptions from the goal summaries of these demo conversations. So the descriptions are close to the messages, and every absolute number here is optimistic.
  • 54 conversations from one demo workload. The cases are few, and the families are synthetic.
  • Qwen3-4B saw one prompt wording and one option order. Another order can give another result.
  • The thresholds tuned on this set are optimistic. The cross-validated numbers are the fair ones, but the set is small.
  • We did not run text-embedding-3-small on the label and description, so this run mixes the effect of the text with the effect of the model.

Source: our own run on 8 October 2026 (scripts/g3_openmodels.py), on the cases of the upstream Strata study at commit ada801f. Models from Hugging Face: mlx-community/Qwen3-4B-Instruct-2507-4bit, sentence-transformers/all-MiniLM-L6-v2 and BAAI/bge-small-en-v1.5. This is research evidence, not live telemetry and not a product feature.

Continue the investigation.