Guessing the task from a chat's first message

Operant·

When a new chat starts, all you have is its first message.

To send it to the right model and instructions, you have to guess the kind of task. A wrong guess gives your user the steps for a different task, and a wrong answer.

Operant sorts your past chats into patterns: kinds of task that repeat. Each pattern can get its own cheaper model and its own instructions. This research asks whether a model can name your pattern from the first message alone. It also has to say "none" when nothing fits. This is research, and it isn't part of the product yet.

Saying "none" is the safe mistake. Your chat just stays on its usual model. A wrong pattern is worse, because it sends the chat to instructions written for another task.

One case tells the story

In the study, seven chats were sanctions screening. That task looks a lot like the triage of suspicious-activity reports, but it has no pattern of its own. So the right answer for all seven was "none".

One small model, Jev, put all seven into suspicious-activity triage, a different task. The other small model, Kev-9B, said "none" in most cases like these. The rest of this post shows how that played out across every case. The screenshots come from the live console, with sample values from a demo workspace.

The patterns in a sample workspace
Four named patterns and a long tail of one-off goals (sample values).

What a pattern looks like

Each pattern has a name, a description and the chats that belong to it. The members tab shows the goal of each of those chats.

The chats in one pattern
Each chat shows a short summary of its goal (sample values).

A first message isn't a tidy goal summary. It's often short, and it often holds an account number or a ticket ID. That's what makes this hard.

The cases

The study used 54 chats from a demo workload:

  • 38 chats belonged to a family with a pattern. The matcher saw four patterns, each with its name and description, plus a "none" option.
  • 16 chats should get "none". Seven were the sanctions-screening chats. Nine were one-off tasks.

The study then asked 38 more questions. For each family chat, it hid that family's pattern, so the right answer became "none". That made 92 cases in total.

The three matchers

  • Jev is a hosted model from TypeSafe, built to pick between options. It gives a probability for each one.
  • Kev-9B is an open version of Jev on Hugging Face. It ran locally with MLX.
  • A similarity match compares the first message with the average of each pattern's goal summaries. It picks the closest pattern when the similarity is 0.5 or more.

Results

Here's the share of the 92 cases each matcher got right:

  • Jev, top option: 85.9%.
  • Jev, only when its top probability was 0.9 or more: 93.5%. We chose that floor after seeing the results.
  • Kev-9B, top option: 87.0%.
  • Similarity match at 0.5: 55.4%.
  • Similarity match at the best setting for this set, 0.56: 66.3%.

The similarity match found the right pattern for only 21.1% of the family chats. Both small models did much better.

Two ways to be wrong

Jev and Kev-9B scored almost the same overall. But they were wrong in opposite ways:

  • Jev found 92.1% of the family chats, but it got only 37.5% of the "none" cases right.
  • Kev-9B found 71.1% of the family chats, and it got 93.8% of the "none" cases right. When it was wrong, it usually said "none".

If you're picking a model, Kev-9B's mistake is the safer one. A floor on the probability helps Jev, though. At a floor of 0.9, it got 87.5% of the "none" cases right. It still found 89.5% of the family chats. A real floor has to come from live traffic, not from this small set.

Open models on the same cases

We then ran three open models from Hugging Face on one Mac, on the same 92 cases:

  • Qwen3-4B, an open instruct model, got 80.4% right. Like Jev, it sent all seven sanctions-screening chats to suspicious-activity triage.
  • Two small open similarity models compared the first message with each pattern's name and description. They got 92.4% and 90.2% right.

The first similarity match used goal summaries. These two read the same text as the small models did. We changed the model and the text at the same time, so the next run has to separate the two. The full table is on the open models results page.

Over all 92 cases, Kev-9B got 87.0% right, and when it was wrong it usually said "none" (full table).

These numbers are optimistic. Operant wrote the pattern descriptions from the same demo chats, and the set is small. Jev's 0.9 floor was also picked after the fact. Every number was recalculated from the study's committed row files (full table).

So a small model could sort most of your new chats from the first message alone.

See every case on the results page.