Article chapter 01 of 08
Start with one state disagreement
From Releasing a multi-portal platform without losing state between components
Start with a specific disagreement and leave the theory about its cause out of the description. For example: a user completes an action, receives confirmation, and an authorised staff member still sees the item as pending. That gives you two roles, two views and one business state. It is narrow enough to reproduce.
Avoid beginning with a broad ticket such as "fix portal synchronisation". That ticket invites a tour of the codebase and encourages several plausible fixes at once. It also hides the question a reviewer will eventually ask: which observable behaviour changed?
Capture the disagreement as a short trace:
- Record the starting identity, role and account or tenant boundary.
- Record the action taken and the data entered.
- Note the immediate response, including any identifier returned.
- Inspect the state shown in each affected portal.
- Inspect queued work, integration records and error state.
- Repeat with the smallest variation that may change the outcome.
Keep timestamps in one declared timezone and retain the raw values where possible. A sequence can look impossible when a browser, application server and external service render the same instant differently. The investigation should distinguish a genuinely late event from a display conversion error.
Keep the trace as a test case even if the first suspected cause is wrong. Without it, each handoff can trigger another attempt to recreate the failure from a different slice of the workflow.