InsightsAI engineering

Evals before demos: how we take AI features to production

A demo shows the best cases. Production sees all of them. An eval suite is how you know which one you have built.

Otorithm Engineering · 6 min read

Most AI features are approved on the strength of a demo. Someone picks a handful of examples, the model handles them well, and the room agrees to build. Then production arrives with the long tail: ambiguous inputs, missing data, users who phrase things in ways nobody expected. Quality turns out to be a distribution, and the demo showed only its best end.

We work the other way round. Before building an AI feature, we agree how it will be tested and what score counts as good enough. The feature is done when it passes, not when it impresses.

Start with the test set

Collect real cases from the workflow the feature will serve: a few hundred is often enough to start. Include the awkward ones on purpose. For each case, write down what a correct result looks like. Then agree the pass mark with the people who own the outcome, before any code exists. This conversation often changes the scope, which is cheaper now than later.

Grade automatically, check by hand

Use deterministic checks wherever the output allows it: schema validity, exact fields, required citations. For open-ended text, use rubric-based grading by a model, calibrated against human ratings on a sample so you know how far to trust it. Keep a regular human review of a random sample, because graders drift too.

Make every change earn its place

Prompts, retrieval settings, model versions and code all change behaviour. Run the eval suite in CI on every change, and treat a regression the way you would treat a failing test: it blocks the merge until someone decides it is acceptable. This is what lets a team improve an AI feature quickly without fear.

Watch production

  • Sample live traffic into a review queue and grade it with the same rubric.
  • Track quality, latency and cost per task side by side.
  • Turn every new failure into a new eval case, so it cannot quietly come back.

Decide with numbers

Evals turn model choice into an engineering decision. If a smaller, cheaper model passes the easy cases at the same rate as a larger one, route those cases to it and keep the larger model for the hard ones. If a new model release scores better on your suite, switching is a measured decision rather than a hope.

A good eval suite outlives any single model. It encodes what your business means by a correct answer, and that is the asset worth building first.