Inside the eval gate

Operant·

Operant can move a repeated task to a cheaper model. It does this only after a gate. The gate replays held-out work on the cheaper model with a learned skill. A judge scores the result against your own past answers. Then a person ratifies the change.

This post explains each step of the gate. The screenshots come from the live console. They show sample values from a demo workspace.

The gate in the live console, in one minute (sample values).

A learned rule starts in shadow

Cortex groups your conversations into patterns by goal. Scribe learns a skill for a pattern. Then Operant proposes a learned rule: this pattern goes to this target model with this skill version.

A new learned rule starts in shadow. In shadow the rule changes no traffic. The gate exists for one step only: the step from shadow to active.

A learned rule on router/ops-assistant
A learned rule after its gate passed: Gate 95%, 3 exemplars, judge claude-opus-4-8 (sample values).

The gate grades on held-out work

Before Scribe writes a file, it saves a split of the pattern into a learn set and a holdout set. By default, one quarter of the conversations go into the holdout. The skill learns only from the learn set.

The gate grades the skill on the holdout set. It never grades the skill on the conversations that taught it. It takes the costliest conversations first, because routing must hold on that part of the traffic. A run uses up to 5 conversations by default.

The provenance of a learned skill
Eight conversations taught this skill. A separate holdout stays for the gate (sample values).

Some patterns have too little captured text. For these, the gate makes one-turn tasks from the goal summaries. It marks each of those scores as synthetic, so you can see the weaker evidence.

Replay against your own baseline

For each conversation in the set, the gate replays the user turns two times:

  • The baseline arm uses the model that the conversation used most.
  • The candidate arm uses the target model with the skill.

The baseline comes from each conversation. So the gate compares the candidate with what your agent did, not with one fixed model.

A judge scores goal completion

An LLM judge reads the same user turns with the two sets of answers. It gives a parity score:

ParityMeaning
1.0The candidate did as well as the baseline, or better.
0.5The candidate did part of the goal.
0.0The candidate did not do the goal.

The judge scores goal completion, correctness and completeness. It does not score style or length. It also writes one sentence that tells why. The default judge model is claude-opus-4-8, and the console shows it on each rule.

The gate passes when the mean parity is 0.8 or more. That is the default bar. The console shows the result on the rule, for example "Gate 95%".

A person ratifies

A passed gate moves no traffic. A person must ratify the rule. After that, the rule sends the pattern to the target model, and the skill goes into the system prompt of each routed call.

The analytics of a pattern also show this path. When a pattern has a learned skill, the finding says: "Approve the skill, run the eval gate, and ratify the route."

A routing finding on a pattern
The finding that tells you to run the eval gate and ratify the route (sample values).

Change anything, run the gate again

The gate approves one set of three things: the target model, the skill version and the compression settings. If you edit one of them, Operant clears the report. Then the gate must run again before you can ratify the rule.

An active rule cannot change its target. First you demote it to shadow. Then you edit it, and the gate runs again.

Demote in one click

The button "Demote to shadow" sends the pattern back to the model it used before, at once. Demote has no gate. If you have a doubt, you can stop the change in one click.

What the gate does not see

Readers on Reddit named three limits of the gate. They are correct, and each one is a limit of the gate today.

It measures agreement, not a correct outcome. The judge compares the candidate with the past answers of your agent. If a past answer was wrong, a candidate with the same answer still gets a good score. So the person who ratifies reads replay scores. Nobody grades the real outcome of each case.

A replay sees only old inputs. The holdout comes from past conversations. So the gate does not see drift: new kinds of requests that come after the switch. The gate keeps no live traffic on the old model. A next step is to keep a small share of live traffic on the old model for some time after a switch. Then you compare the outcome scores. Until then, "Demote to shadow" is the fast way back.

A good mean can hide a bad case. The bar is on the mean parity, so one costly failure can hide in a good mean. The default judge also reads each answer as text. The gate has no separate format check, for example a check that each required field is there. Next steps are format checks and a second bar on the worst 10 percent of cases.

For the step before the gate, read How Scribe learns a skill from your own traces. For the judge, read Can a small decision model grade an agent's outcome?.