You want your model bill down. There are many ways to do it, from a clearer instruction to training your own model.
Do them in the wrong order and it costs you either way. Train too early, and you spend weeks on a task a clearer instruction would have fixed. Never move, and you pay top rates forever for work a smaller model could do.
Operant follows a ladder with three steps. Each step is earned by the one below it.
| Step | What it gives you | Status |
|---|---|---|
| 1. Find the work that repeats | The few tasks that drive most of your spend | Shipped |
| 2. Prove a cheaper setup holds | A cheaper model that passed a test on your own work | Shipped |
| 3. Train your own model | Weights you run yourself, for one proven task | Not built |
Step 1: find the work that repeats
Before you change anything, you need to know what the work is. Operant captures every call, groups calls into conversations, and sorts conversations into named kinds of task. Spend piles up in a few of them. On our own development setup, the top three kinds of task were 46% of spend. That's our own number, not a market study, but the pile-up is the point. You can't cut the cost of work you can't name.
Step 2: prove a cheaper setup holds
For the biggest slice, Operant learns a skill from that task's own successful chats. A skill is short instructions that cover the tools, the judgment and the output format the task needs. Then Operant runs chats the skill never saw through a cheaper model with the skill. It scores the answers against the top model's own answers.
Cheap fixes come first here too. A tighter instruction or a fixed output format often moves the score before any model switch does. Whatever the fix, it has to pass that same test, or it doesn't ship.
Step 3: train your own model
The top step is training: turning a proven slice into weights you run yourself. We haven't built it, and we say so plainly. Today it's on the roadmap, not in the product.
What steps one and two hand it is exactly what training needs. That's the grouped and priced data, the slice worth the money and a pass mark the test already used. Training before those exist is how teams burn whole quarters.
Why the order is the product
Every step shares one check: a test on chats the change never saw, then a person's approval. That makes the ladder safe to climb and safe to come down. Turning a change off needs no test, and the top model always stays as the backup.
On our own development setup, step one found that just three kinds of task drove almost half of all spend. Steps one and two are shipped, and step three isn't built yet.
So every step you take has already been proven by the step below it.
See the test that guards each step in How we prove a cheaper model can do the job.
