Blog
What applied AI means in my work
Applied AI starts with a piece of work that already exists. I want to understand the decision, the information available, the cost of delay and what happens when the answer is wrong before choosing a model or building an interface.

A request for an AI feature often arrives as a capability: summarise documents, classify messages, answer questions, generate a response. I first locate the decision or action that would use the output. “Summarise this document” could mean helping a reader decide whether to open it, extracting obligations for a project plan, or preparing a briefing for someone who has not seen the source. Each job needs different information. A fluent paragraph may be fine for orientation and useless for obligation tracking, where dates, responsible parties and direct links back to the source matter. I therefore describe the work without naming a model. Who receives the input? What do they do with it now? Which part consumes time? Where does judgement enter? What record must remain after the work is complete? If the workflow cannot be explained in ordinary operational terms, adding AI usually makes the ambiguity harder to see. A model cannot repair a workflow that lacks the information needed for its decision. It may produce a plausible answer anyway, which makes source quality an implementation issue rather than a background concern. The source inventory should be concrete. List the documents, database fields, messages or images available at the point of work. Record who can access each source, how current it is, and what happens when sources disagree. A policy file with no effective date is not equivalent to a current policy. A customer record copied into a prompt may be incomplete or outside the user's permission.
For retrieval-based work, I want to know whether the system can return the exact passage supporting an answer. For classification, I want labelled examples that reflect the cases the team actually handles, including ambiguous and incomplete ones. For generation, I want to know which facts must come from a controlled system rather than from the model's general knowledge. This source work can reveal that conventional software is enough. If a decision follows stable rules over structured fields, a query or rules engine may be easier to test and explain. I compare that option with a model on the same cases before adding the model to the workflow. One accuracy figure hides which cases failed. A wrong category on an internal queue may create a small delay. A wrong instruction sent to a customer can create confusion, rework or a commitment the organisation did not intend to make. I set release checks for those cases separately. Before building, I separate outcomes by consequence. Which errors can be corrected before anyone relies on them? Which can be reversed afterwards? Which could expose private information, move money, change access or affect a person's rights? The workflow should add review and approval where the consequence demands it.
Human review works only when the reviewer has the source, enough time and an interface that makes uncertainty visible. Asking someone to approve a long generated answer without showing the source encourages a quick visual scan. It may add delay without catching many errors. The useful measure is task success after the full review process. If the feature saves two minutes of drafting and adds three minutes of checking, the impressive generation speed is irrelevant to the operating decision. Some work is valuable because it is correct. Some is valuable because it arrives before a deadline. Applied AI design has to account for both. A task with low volume and high consequence may suit careful retrieval, slower processing and explicit approval. A high-volume intake queue may need fast classification, confidence thresholds and a route for uncertain items. A live conversation has tighter latency constraints, but speed should not cause the system to take an irreversible action before the user confirms what it understood. Volume also affects cost and operations. A model call that looks cheap in a demonstration can multiply across retries, long prompts, repeated source retrieval and reviewer time. The calculation should include the whole task: preparing the input, processing it, checking it, correcting it and keeping the resulting record.
I prefer to test this with a sample of real task shapes rather than an average document or ideal prompt. Long inputs, missing fields, duplicated requests and mixed-language material often decide whether a workflow remains usable. An AI system does not have to complete the whole job to be useful. It might extract candidate fields for confirmation, rank a queue, identify passages that deserve attention, or draft a response that remains visibly unapproved. With one action, the team can name the input, the person who receives the output and what happens next. It also gives the evaluation a concrete target. “Did it extract the correct invoice date?” is easier to test than “Was the assistant helpful?” Accepted and rejected cases should determine whether the feature takes on more work. Repeated correction of the same field points to a system change or a task the model may not suit. If reviewers accept outputs without reading the source, the interface may be encouraging unsafe use. Starting small is also a way to learn the workflow. The first release often exposes exceptions that interviews and process diagrams missed.
I want a set of representative cases before choosing the final model or designing a conversational screen. The set should include routine examples, difficult examples, incomplete inputs and cases where the correct action is to decline or ask for more information. For each case, define what an acceptable result contains and what would make it unsafe or unusable. Keep the source and expected outcome beside the model output. When a prompt, model or retrieval method changes, run the same cases again and compare the effect on the task, not just the prose. Then test the workflow around the model. Can the user inspect the source? Can they correct a result? Is the correction stored? Can an operator find failed runs? Does a retry duplicate an action? These questions decide whether the feature can live inside real work. For one existing task, I write down the input, the decision being supported, the consequence of an error, who checks the result and where the accepted result is recorded. I use that description to choose the evaluation cases before comparing models.
Discussion
Continue the thinking.
Comments are public and hosted in an open-source GitHub Discussions repository.
Loading comments connects your browser to GitHub. A GitHub account is required to post.