20 February 2024 / Applied AI

How I decide whether an AI idea deserves a build

A convincing demo is easy to mistake for a product opportunity. I first ask who would notice a wrong result, what they could check and whether the time saved survives the review work the feature creates.

A designer drawing a building by hand at a desk.
Photo: Ryan Ancill (opens in a new tab)

A demonstration usually shows the strongest moment: a document goes in and a polished answer comes out. A product has to deal with everything around that moment. It needs to select the right document, check access, handle missing information, present the answer to the right person, record any correction and recover when a downstream step fails. I start by writing the task as a sequence with a clear beginning and end. If the feature drafts a response, the task does not end when text appears. It ends when the response has been checked, approved, sent and attached to the correct record. If the feature extracts fields, the task ends when accepted values reach their destination without overwriting better data. This immediately changes the build estimate. The model call may be a small part of the work. Identity, source retrieval, user interface, audit records, exception queues and support tools can take more effort than the generation step. At this point I can assess the complete task, the expected outcome and who would operate it. The isolated model output is only one part of that assessment. Every proposed output needs a realistic error detector. “A human will review it” is incomplete. Which human? What evidence will they see? Do they know enough to recognise a subtle error? How much time can they spend on each case?

For a summary, the reviewer may need direct links to the passages behind factual statements. For data extraction, they may need the source image beside each proposed value. For classification, they may need the category definition and a route for uncertain cases. The interface should support the act of checking, rather than putting an approve button under a wall of generated text. Some errors surface only after the result leaves the review screen. A misrouted request may be noticed by the team receiving it. A wrong account update may be noticed by support. Those people are part of the control design, and the system needs a way for their correction to reach the original record. If nobody can detect a wrong result before the consequence becomes expensive, I narrow the feature or keep it away from that action. Time saved in generation is easy to show. Review time is spread across people and often disappears from the calculation. It includes opening sources, checking claims, correcting format, resolving uncertainty and occasionally rebuilding the output from scratch. A fair comparison uses the current task as the baseline. Measure active handling time and waiting time separately. Then run representative cases through the proposed process and record:

  • preparation time before the model can run;
  • model and retrieval time;
  • reviewer time, including source checks;
  • correction and escalation time;
  • repeated work caused by failed or duplicate runs.

The sample needs ordinary and difficult cases. Testing only short, clean inputs produces a number that will not survive normal use. Review may also change as users learn where the system tends to fail, so early observations should be treated as provisional. If the new workflow moves work from one team to another, that still counts. A feature can make one screen faster while increasing support and reconciliation work elsewhere. AI ideas often point towards a real problem but not the cheapest useful response. Poor search, missing fields, unclear status, duplicated data and weak templates can all make a generative feature look more necessary than it is. I compare the proposed build with at least one conventional change. Could a better query find the relevant records? Could a required field remove the need to infer a value? Could a template with approved clauses handle most responses? Could a rules engine route the cases that follow stable conditions, leaving only ambiguous items for review? The comparison should use the same task cases and acceptance criteria. It is unfair to demand perfect performance from a simple rule while accepting an AI answer because it sounds reasonable. It is equally unhelpful to reject AI because it cannot cover rare exceptions when the existing process cannot either.

The design may combine deterministic code with a model. Code can enforce permissions, amounts and state changes while the model handles unstructured input. I record which component is responsible for each decision so a failure can be traced to the right part. In a first release, drafts, rankings and extracted fields should remain proposals that the user can inspect before an external action occurs. Reversibility also needs technical support. Keep the original input. Store the proposed output separately from the accepted value. Record who approved a change and when. Use idempotency controls where retries could repeat an action. Give operators a way to locate runs that stopped between steps. This state is useful during evaluation. It lets the team compare proposals with accepted results and inspect recurring corrections. Without it, feedback becomes a collection of impressions: people say the feature is “pretty good” or “unreliable”, but nobody can isolate the cases behind either view. For higher-consequence work, use a limited group and explicit boundaries. The release should state what the feature may do, what it cannot do, and where uncertain items go. Expanding access is a later decision supported by the first run's evidence.

I want a short decision record before implementation starts. It should name the user, task, current baseline, proposed intervention, required sources, error consequences and review method. It should also state what result would stop or reshape the build. The evaluation set belongs in that record. Gather cases from the real task shape, remove private detail where the test environment requires it, and define acceptable outcomes. Include inputs with missing evidence and cases where the system should refuse to proceed. Then decide how results will be compared. Depending on the task, that may include field accuracy, correct routing, unsupported claims, reviewer corrections, total handling time and unresolved cases. Avoid combining all of these into one score too early. A small number can hide a failure that matters operationally. I start with a one-page task and evaluation brief, then run representative cases manually. The result needs to include review time, corrections and failures. That gives me enough information to estimate the build or drop the idea before implementation begins.

Continue the thinking.

Comments are public and hosted in an open-source GitHub Discussions repository.

Loading comments connects your browser to GitHub. A GitHub account is required to post.

All blogs