Article chapter 08 of 08
Plan reconciliation and rollback before you release
From Releasing a multi-portal platform without losing state between components
The final review should use the same state model as the implementation. A generic sign-off list can confirm the deployment is healthy and still miss records that an earlier version left stranded.
Before release, I'd check:
- every business fact in the workflow has one named owner
- role checks cover finding, reading, transitioning and reversing
- duplicated and delayed integration events are safe
- uncertain operations go into reconciliation instead of disappearing
- identifiers connect user reports, local records and external events
- the original composite failure patterns have regression tests
- each portal's configuration points at the services it's meant to
- queues and exception records are visible to an operator
- rollback doesn't leave the data model ahead of the running code
- someone is assigned to review exceptions after deployment
Work out what rollback means for state that's already changed. Reverting the code won't undo an external action or bring back an earlier permission decision. The release plan might need a forward repair, a temporary pause on one transition, or a script that only reports the affected records so someone can handle them by hand. Test that script on a copy, or in report-only mode, before it's allowed to write anything.
After deployment, run a small set of end-to-end traces with test records you can identify. Go through the reconciliation queue, and only compare event counts where those counts have a defined relationship to each other. I'd keep the release open until the team can explain every test operation and any exception it produced.
For the workflow you've just released, check that every authorised role sees a state that matches the owning system. Then force each supported failure path and confirm it leaves enough evidence for the named recovery action to work.
Picture the applicant again, submitting in one portal while the operator sees an older version in another. Releasing a multi-portal platform without losing state between components mostly comes down to each business fact having one owner, every other component having a defined way to find out it changed, and uncertain integration work staying visible until it's reconciled. When the next disagreement turns up, I'd write it down as a short trace first, find the one contract that's failing, and hand that to the agent or reviewer as a bounded work order with tests in both affected roles.