Article chapter 05 of 08
Put exceptions and corrections on the map
The normal path only covers part of what you're replacing. Support effort tends to pile up around incomplete requests, duplicate events, changed decisions, services that are down and records that disagree with each other. Put those cases into the process with named recovery paths, instead of leaving them on a miscellaneous edge-case list.
Ask about categories of exception rather than stories, because stories can expose private details. What can be missing when a request arrives? Which values can conflict? Which actions can happen twice? What can change after approval? Which external responses can be late or never turn up? You can then check those categories against anonymised records or policy material.
For each exception, write down detection and recovery separately. Detection might come from validation, a reconciliation job, someone noticing an inconsistency or the requester following up. Recovery says who can fix it, which state the work goes back to, and whether later steps have to be reversed or run again. A generic "failed" status usually hides all of that.
Correction rules need a close look. If an approved value changes, does the product keep the original, create a revision or overwrite it? Who's allowed to change it? Does it need approving again? Which downstream systems have already received the earlier value? Those answers shape the data model, audit history, permissions and integration design all at once.
Duplicates are another good test. A request might get submitted twice, a payment notification might be delivered again, or someone might retry after a slow response. The team needs a way to recognise the same business event without blocking legitimate repeat activity. That usually means an identifier, a time window or a comparison rule that's spelled out more precisely than "ignore duplicates".
Don't forget abandonment and cancellation. Work sometimes stops because the requester withdraws, the information never arrives or the task stops applying. The new system should be able to tell an intentional stop apart from an item that just fell out of a queue. Write down who can close it, what reason they have to give and whether it can be reopened.
Go through the exceptions with the people who actually fix problems, as well as the owners of the main process. They know where the evidence runs out and which changes are hard to undo.