The question we are asked most about a language-model feature is "how accurate is it". The honest answer, for most teams, is that nobody knows, because nobody wrote down what a correct answer looks like before the model started producing them. A demo that impresses on ten questions says nothing about the next thousand.
So the first thing we build for an LLM feature is not a prompt or a pipeline. It is an evaluation set: real questions, agreed answers, and a way to score a new version against them in minutes. Everything else is measured against it.
What an evaluation set is
An evaluation set is a list of inputs the feature will actually receive, each paired with what a good output looks like and how to judge it. For a document assistant that is a question, the passages that answer it, and the answer a subject expert would accept. For a classifier it is an input and its label. For a summariser it is a source and the facts a summary must contain.
It is not a test of the model in general. It is a test of your feature on your data, which is the only thing that predicts how the feature behaves once it is live.
Why it comes first
Three reasons, in the order clients tend to feel them.
It turns "is it good enough" into a number that moves. A prompt change, a model upgrade or a new retrieval index either raises the score or lowers it, and the discussion is about the cases that changed rather than about impressions.
It stops regressions. Model providers change models under the same name. A retrieval tweak that fixes one question breaks three others. Without a set to re-run, those breaks reach users.
It makes the acceptance criteria real. A contract that says "the assistant answers accurately" cannot be signed off. One that says "at least 90 percent of the evaluation set scores correct, and no answer cites a document it did not retrieve" can.
How we build one in a week
The set does not have to be large to be useful. Fifty to two hundred cases cover most features; a few thousand is for a mature product with traffic to sample from.
Day one, collect real inputs. Support tickets, search logs, the questions the sales team is asked, the emails a back office answers. Real phrasing matters: people misspell, abbreviate and ask two things at once. If nothing exists yet, the people who will use the feature write twenty questions each, in their own words.
Day two, group and stratify. Sort the inputs by type: simple lookups, questions that need two documents, questions the system should refuse, questions with no answer in the data, adversarial inputs. Make sure each group is represented, because the average hides the group that fails.
Days three and four, write the gold answers. A subject expert writes or approves the answer for each case, and marks which source passages support it. This is the slow part and the part that cannot be delegated to the engineers. It is also where the client learns what the data does not contain.
Day five, decide how each case is scored. Some cases have an exact answer and can be checked by string match or a number.
The five numbers we report
Every sprint, the feature is run against the full set and five numbers go on the demo slide.
- Correctness. The share of cases scored correct by the rubric, overall and per group.
- Faithfulness. For retrieval-based features, the share of answers whose claims are supported by the passages the system actually retrieved. An answer can be correct and unfaithful, which means it will be wrong on a different day.
- Refusal behaviour. The share of should-refuse cases correctly refused, and of answerable cases wrongly refused. Both matter; an assistant that refuses everything scores perfectly on safety and is useless.
- Latency. Median and 95th percentile, end to end, because a feature that is right in eight seconds will not be used.
- Cost per query. Tokens in and out at current prices, so a model upgrade is a decision with a bill attached.
A sixth line lists the cases that changed since last sprint, with the old and new outputs. That list is what the review meeting reads.
What the set changes about the build
Once the set exists, the engineering choices become experiments instead of opinions. Chunk size, the number of passages retrieved, the reranker, the prompt wording, the model, and whether to fine-tune at all are each tried against the set and kept only if the numbers improve. The companion article on retrieval versus fine-tuning describes those trade-offs; the evaluation set is what makes it possible to choose between them for your documents rather than in general.
It also changes the conversation with the model provider. When a provider deprecates a model, the migration is a re-run of the set and a look at the changed cases, not a week of manual testing.
Keeping it honest after launch
The set is a starting point, not a monument. After launch, sample live traffic every week, add the cases the feature got wrong, and add the cases users asked that nobody predicted. Retire cases that no longer represent real use. Keep the gold answers under review by the same experts, because the business changes and so do the correct answers.
Two rules keep the set honest. Never tune the feature on the exact cases in the set and then report the score on them; hold part of the set back. And never let the engineers write the gold answers alone.
Where to start
If you have a feature in mind, the evaluation set is a one-week engagement on its own, and it is worth doing before committing to a build. You finish the week knowing what your data can answer, what it cannot, and what "good" will be measured against. How the full build runs is on the generative AI and LLM solutions page; send us a description of the feature and the documents behind it and we will scope that week.


