27 February 2025 / Data systems / 8 chapters

A stopped run exposes the missing contract

From Designing resumable large-data pipelines

An import has processed part of a file when the worker stops. Some records are in the destination, some were rejected, and the rest may never have been read. The source file still exists, but rerunning it from the top could repeat emails, duplicate ledger entries, overwrite newer data or produce a different set of errors.

At that point, recovery depends on one answer: which unit of work can safely run next? The pipeline needs progress records that match its side effects. A log line saying processed 48% is useful for display but weak for recovery. It does not identify accepted records, committed batches or external actions.

Before changing code, write down the job's recovery contract:

  • What is one independently identifiable work item?
  • What counts as accepted?
  • Which writes happen before acceptance is recorded?
  • Can any write occur outside the main database transaction?
  • What evidence proves that an item already completed?
  • What should happen to a record that repeatedly fails?
  • Who can start, pause, cancel and resume a run?

These answers determine the state model. Add progress displays, retry loops and longer timeouts only after the recovery contract is explicit. Otherwise those controls repeat work without resolving which effects already happened.

Design for interruption after any durable effect. A process may stop between an API response and a local status update, between two database commits, or after a file has been moved but before the run summary is written. Recovery has to handle those gaps as part of the normal state model.

All articles