27 February 2025 / Data systems / 8 chapters

Resume through a controlled scheduler

From Designing resumable large-data pipelines

Resumption should be a scheduler operation, not a worker improvising after a crash. The scheduler knows the run's requested scope, current policy, outstanding leases and stop condition. Workers should claim eligible items and report outcomes without deciding that an old run is safe to restart.

Use leases for claimed work. A lease has an owner, token and expiry time. A worker may complete an item only while presenting the current token. If the worker disappears, the scheduler can return the item to an eligible state after expiry. A heartbeat may extend long operations, but set a maximum sensible duration so a stuck process does not hold work indefinitely.

Pausing has two stages. First, stop issuing new claims. Then allow active work to finish or reach an operation-specific safe boundary. The run becomes paused only when no active lease can still produce an unrecorded effect. A force stop may be necessary, but display it as a different action because it can create uncertain items that require reconciliation.

Resume should use the original run definition unless a person deliberately creates a revised run. Parser code, mapping rules and reference data can change while a job is paused. Store their versions and decide whether the old run remains executable. Quietly resuming under new rules makes one run internally inconsistent.

Cancellation also needs a defined meaning. It normally stops future work; it does not reverse accepted records. If reversal is required, model it as a separate compensating job with its own item states and evidence. Calling that behaviour cancel hides consequential work behind an ordinary control.

All articles