19 February 2026 / Software delivery / 8 chapters

Treat each integration as a state machine

From Releasing a multi-portal platform without losing state between components

An integration call is rarely one clean step. The application prepares a request, sends it, receives or fails to receive a response, stores local state, and may later accept a webhook or run a reconciliation job. Any boundary can fail after the external system has acted.

Model the local integration record with explicit states appropriate to the operation. Names vary, but the distinctions usually include work that is waiting, in progress, accepted, rejected, uncertain and reconciled. "Uncertain" matters when a request times out and the caller cannot tell whether the remote side completed it. Retrying blindly can duplicate the action.

Use an idempotency key or stable operation identifier where the external contract supports one. Store that identifier before the request leaves the application. A retry should continue the same operation rather than create a new one. Where the provider offers no idempotency support, the recovery path may need a lookup or a human decision instead of an automatic retry.

Treat each webhook as an external event that still has to pass verification and local transition rules. Verify its origin using the mechanism defined by the provider, retain the external event identifier, and make repeated delivery safe. Process the event against the current local state. An old event arriving late should not overwrite a newer confirmed state.

Add reconciliation for records that remain uncertain or in progress beyond the expected period. Reconciliation should compare identifiers and relevant facts, record what it found, and either complete the transition or put the item into a visible exception queue. A scheduled job that silently scans and skips errors creates the appearance of reliability while leaving support with no case to inspect.

All articles