Blog
Write the evaluation set while the AI feature is still moving
Teams often postpone evaluation until the prompt and interface feel finished. I prefer collecting difficult real cases during development because they reveal whether a change improves the underlying task or simply makes the latest demonstration look cleaner.

Teams often postpone evaluation until the prompt and interface feel finished. I prefer collecting difficult real cases during development because they reveal whether a change improves the underlying task or simply makes the latest demonstration look cleaner. The useful cases are usually already present. They appear when a developer changes the prompt after an awkward output, when a subject-matter reviewer corrects an answer or when a test user has to explain what the feature misunderstood. If those examples stay in chat threads and temporary notes, the team loses them. The next improvement can quietly reintroduce the same problem. A case is worth recording when it changes the design, exposes uncertainty or forces someone to decide what an acceptable result looks like. Waiting for a formal evaluation phase makes that context hard to recover. I keep the first record small. It needs the input, the relevant source material, the observed output and a plain description of the problem. If a reviewer supplies a corrected result, that belongs with the case. So do any conditions that matter, such as the user role, rule version or document date. The aim is to preserve what made the case difficult. A vague note such as "answer was poor" will not help later. "The answer used a superseded policy even though the current policy was present" gives the team a behaviour it can test. "The assistant asked for information already provided in the form" is similarly checkable.
Private material needs care at this point. Evaluation fixtures should use approved, redacted or synthetic inputs where possible. If a real case must remain restricted, store it in a controlled test environment and record the access requirement. Copying private content into a convenient shared spreadsheet creates a separate problem. Once a case is visible, it is tempting to adjust the prompt until that single output reads well. The team should first describe what success means in a way that applies beyond the example. Some checks can be exact. A classification must come from an allowed set. A calculated value must match the versioned rules. A citation must point to a source included in the provided material. An action must remain within the user's permissions. Other outputs need human judgement. A summary may be accurate but omit the point needed for the next decision. A draft may contain all required facts yet use a tone that is unsuitable for the recipient. In those cases, a short scoring guide helps reviewers apply the same criteria. It can state the required information, unacceptable errors and what deserves escalation rather than a confident answer. I also record when several answers could be acceptable. Treating one preferred sentence as the only correct response rewards imitation and can penalise a useful result that is phrased differently.
The team will look at development cases while changing the feature. That is expected. Those cases guide prompt changes, retrieval work, interface decisions and tool behaviour. Passing them shows that known problems have been addressed. A smaller held-back set gives a different signal. Developers do not tune directly against it, so it is more likely to show whether an improvement transfers to similar work. The held-back set still needs review as the product changes. A case can become obsolete when a policy, workflow or feature boundary changes. The split does not need to be elaborate at the start. What matters is recording which cases influenced the change and which were used afterwards to check it. Without that distinction, a reported pass rate can mostly describe examples the team has already rehearsed. Repeated or near-duplicate cases also need attention. Ten versions of the same easy request can make the result look stable while one uncommon, consequential case remains untested. I group related cases and look at coverage by task, failure type and consequence rather than relying on one total. Evaluation results are difficult to interpret when the configuration moves during the run. Each result should identify the model, system instructions, retrieval settings, tool versions, rule versions and any sampling controls that affect the output.
External dependencies matter as well. A search index may change between runs. A connected service may return different data. A tool timeout can turn a correct plan into an incomplete result. Where possible, I use fixtures or snapshots for dependencies that need repeatable testing. Separate live integration checks can confirm that the current systems still connect properly. The runner should retain the raw output and the check result. A single pass or fail value hides useful information, especially when an automated check is wrong or a reviewer disagrees. Keeping the evidence makes it possible to inspect the result without trying to recreate an older configuration. An overall score is useful for spotting a large regression, but it rarely tells the team what to change. I want to see which task failed, how it failed and what would happen if that output reached the workflow. A formatting error may be annoying and easy to recover from. A confident answer based on the wrong source needs different handling. So does an agent action directed at the wrong record. Grouping failures this way helps the team decide whether to change the prompt, improve retrieval, narrow tool access, add a deterministic check or require review.
Human review effort belongs in the result too. If a change improves automated accuracy while making every answer slower to inspect, the product trade-off has changed. Record the correction or escalation needed, even if the output eventually becomes usable. When a new difficult case appears, add it with the reason it matters. When a case becomes invalid, retire it with an explanation instead of silently deleting it. When a change is proposed, run the relevant cases and keep the result beside the work. A practical pull request or release check can state which evaluation version ran, which configuration it used and which failures remain accepted. The first useful step is even smaller: take the next example that causes someone to retune the feature and turn it into a repeatable fixture before the discussion moves on.
Discussion
Continue the thinking.
Comments are public and hosted in an open-source GitHub Discussions repository.
Loading comments connects your browser to GitHub. A GitHub account is required to post.