Article chapter 08 of 08
Testing what happens when you stop it at awkward moments
A normal successful run tells you very little about whether the pipeline can resume. Recovery tests should interrupt it around each durable boundary and then look at the state it's left in.
Build fixtures for duplicate identifiers, changed checksums, invalid references, dependency timeouts and records that fail after partial preparation. Then force an interruption:
- before an item is claimed
- after the claim but before the first write
- after a destination write but before local acceptance
- between items in a committed batch
- during a pause while workers still hold leases
- after an external request returns but before its response is stored
- while the run summary is being calculated
For each one, resume the run and check the destination records, item states, attempt history and aggregate counts. You're looking for accepted items that weren't repeated, uncertain effects that got reconciled, rejected items that are still visible and pending work that carried on.
Test the operator's side as well. The run page should show its source and configuration versions, the last committed progress, active and expired leases, grouped failures, quarantined items and what action is available next. Someone looking at a bare percentage doesn't have enough to make a recovery decision.
Before release, I'd do a witnessed stop and restart against a production-like destination. Write down the item count before the interruption, the exact boundary you stopped at, the state straight after stopping and the reconciled result after resuming. Can the team explain every item at that point? If not, the pipeline is still counting on a clean run that never gets interrupted.
So back to the import that stopped halfway through the file. Whether you can carry on from there comes down to the pipeline's progress records matching what it actually did: item identities that stay the same between attempts, checkpoints only after durable effects, writes that are safe to repeat and failed records set aside where someone can see them. If you've got a job like this, I'd write its recovery contract first and then stop it at one of the awkward boundaries above, well before anyone needs to decide whether it's safe to run from the top again.