LLM evals for small teams: a practical starting point
Most small teams know they should have LLM evals. Few have them, because the word suggests a research project: benchmarks, dashboards, a platform to buy. The useful version is much smaller. It is a fixed set of real conversations, a few checks you trust, and a rule that nothing ships until it passes.
We run offline evals for two live products, Cabin and Mo, the AI friend inside QuesMo. They share one backend, so every prompt or model change is checked against both. This is the process we would hand to any team starting today.
Key takeaways
- Build the eval set from real failures, not from imagined edge cases.
- Write binary pass/fail checks. Scores out of five hide disagreement.
- Use code for anything code can check. Save the LLM judge for judgement.
- Trust a judge only after it agrees with a person on labelled examples.
- Run the suite on every prompt, model and memory change, before users see it.
Start by reading, not by building
The first step has no tooling at all. Pull 30 to 50 recent conversations and read them. For each one, write a short note on what went wrong, in plain words: “reply too long”, “asked a question it already had the answer to”, “ignored the budget the person gave”.
Then group the notes. After a few dozen, the same handful of failure types keep showing up. That list is your first eval plan. Practitioners call this error analysis, and it is the step most teams skip. Skipping it means you end up measuring what was easy to measure instead of what is actually going wrong.
For Cabin, the checks that matter most follow the product’s promise: short replies with a clear opinion. So some of the first checks are about length, and about whether the reply takes a position or hides behind “it depends”.
Turn each failure into a test case
Each failure type becomes a small set of test cases. A test case is an input, usually a short conversation, plus a way to decide whether the output passed.
Anthropic’s guide to agent evals suggests starting with 20 to 50 tasks drawn from real failures, and that matches our experience. You do not need hundreds on day one. You need the ones that hurt.
A few rules help:
- Keep the conversation that caused the bug. Paraphrasing it into a clean example often removes the thing that tripped the model.
- Include the negative case. If you test that the reply mentions a remembered fact, also test a conversation where it should not.
- Write the expected behaviour down before running the model. Otherwise you drift into grading whatever came out.
Memory needs its own cases. The most useful ones are conversations where a fact changes halfway through, because the right reply depends on the latest state. We wrote about why in designing AI memory that stays accurate.
Pick the cheapest grader that works
There are three kinds of grader, and each has a place.
Code checks
Anything a rule can decide should be a rule. Reply length, banned phrases, valid JSON, whether a required field is present, whether a crisis resource appears when it must. These run in milliseconds, never disagree with themselves and cost nothing.
An LLM judge
Some questions need judgement: did the reply give a clear opinion, did it sound like a person, did it use the remembered fact naturally. For these, a second model reads the conversation and a written rubric, and answers pass or fail with a one-line reason.
Ask for pass or fail, not a score out of five. A 3 and a 4 mean different things to different readers, and to the same judge on different days. A binary answer forces the rubric to say what good looks like.
A person
A person is the reference that everything else is measured against. You do not need people to grade every run. You need them to label enough examples to check the judge.
Calibrate the judge before you trust it
An LLM judge is a model, so it can be wrong in quiet ways. Before relying on one, label a set of outputs by hand, run the judge on the same set, and compare. Look at the two error types separately: failures the judge missed, and passes it wrongly flagged.
When they disagree, read the cases. Usually the rubric is vague, or the judge is missing context the person had. Fix the rubric and check again. Keep a held-out set that you never tune against, so you know the agreement is real.
Recheck the judge when you change the judge’s model. A new model can grade differently even with the same rubric.
Run it on every change
The suite earns its keep when it runs before anything ships. In our setup, a prompt edit, a model swap, a routing change or a memory change all trigger the same run: replay the fixed conversations through the current version and the proposed one, then compare pass rates per check.
Compare per check, not only the total. A change that raises the overall pass rate while breaking one important check is not an improvement. This matters most when you route between LLM providers to save cost, because a cheaper model often fails in narrow, specific ways.
Two habits keep the suite useful over time:
- Every new bug becomes a test case. The suite grows from production, not from a planning meeting.
- Read a sample of transcripts each week. Graders can pass outputs that a person would reject. Reading is how you notice.
What this looks like in a small team
None of this needs a platform. Ours is small too: a fixed set of conversations, a script that replays them, rule checks for what rules can decide, a written rubric for the judgement calls, and results compared per run. It fits the way we already work: a spec, a plan, a check before anything is accepted. More on that loop on our engineering page.
The takeaway is simple. Build the eval set before you need it. The day a new model launches, or a cheaper one tempts you, is a bad day to discover you have no way to compare. If you want help setting this up for your own product, we offer it as an LLM audit, or you can get in touch.
Work with usNeed something like this built? Say hello