Article chapter 07 of 08
Resuming through the scheduler
I'd make resuming something the scheduler does, rather than leaving a worker to improvise after a crash. The scheduler knows the run's requested scope, the current policy, which leases are outstanding and what the stop condition is. Workers should claim eligible items and report outcomes. They shouldn't be the ones deciding that an old run is safe to restart.
Use leases for claimed work. A lease has an owner, a token and an expiry time, and a worker can only complete an item while it holds the current token. If the worker disappears, the scheduler can put the item back up for claiming once the lease expires. A heartbeat can extend the lease for long operations, but set a sensible maximum so a stuck process can't sit on work forever.
Pausing happens in two steps. First, stop handing out new claims. Then let active work finish, or get to whatever safe stopping point that particular operation has. The run should only move to paused once no active lease can still produce an effect that hasn't been recorded. Sometimes you'll need a force stop, but show it as a separate action, because it can leave items in an uncertain state that needs reconciling.
A resume should use the original run definition unless someone deliberately creates a revised run. Parser code, mapping rules and reference data can all change while a job is paused. Store their versions and decide whether the old run can still be executed. If it quietly resumes under new rules, you end up with one run that's inconsistent with itself.
Cancelling needs a defined meaning too. Normally it stops future work and leaves accepted records alone. If you need to reverse things, model that as a separate compensating job with its own item states and evidence. Calling that cancel hides a lot of consequential work behind a button that looks harmless.