27 February 2025 / Data systems / 8 chapters

Test recovery at the awkward boundaries

From Designing resumable large-data pipelines

A normal successful run proves very little about resumption. Recovery tests should interrupt the pipeline around each durable boundary and inspect the resulting state.

Build fixtures for duplicate identifiers, changed checksums, invalid references, dependency timeouts and records that fail after partial preparation. Then force interruption:

  • before an item is claimed;
  • after claim but before the first write;
  • after a destination write but before local acceptance;
  • between items in a committed batch;
  • during pause while workers still hold leases;
  • after an external request returns but before its response is stored;
  • while the run summary is being calculated.

For each case, resume the run and check destination records, item states, attempt history and aggregate counts. Confirm that accepted items were not repeated, uncertain effects were reconciled, rejected items remained visible and pending work continued.

Also test the operator path. The run page should identify its source and configuration versions, last committed progress, active and expired leases, grouped failures, quarantined items and the action available next. A percentage without those details cannot support a recovery decision.

The final release check is a witnessed stop and restart against a production-like destination. Record the item count before interruption, the exact boundary used, the state immediately after stopping and the reconciled result after resumption. If the team cannot explain every item at that point, the pipeline is still relying on a clean uninterrupted run.

All articles