Article chapter 01 of 08
What you need to know when a run stops halfway
Say you've got an import that's partway through a file when the worker stops. Some records are already in the destination, some got rejected, and the rest may never have been read. The source file is still sitting there, so the obvious move is to run it again from the top. The trouble is that doing that could send the same emails twice, duplicate ledger entries, overwrite data someone has changed since, or give you a different set of errors from the first run.
So what you actually need to know is which piece of work is safe to run next. For that, the pipeline has to keep progress records that line up with what it actually did. A log line saying processed 48% is fine for a progress bar, but it's not much use when you're recovering, because it doesn't tell you which records were accepted, which batches were committed or which external calls already went out.
Before I touched any code, I'd write down the job's recovery contract by answering these:
- What's one work item that you can identify on its own?
- What counts as accepted?
- Which writes happen before acceptance gets recorded?
- Can anything get written outside the main database transaction?
- How do you tell that an item has already completed?
- What happens to a record that keeps failing?
- Who's allowed to start, pause, cancel and resume a run?
Your answers decide what the state model looks like. I'd leave progress displays, retry loops and longer timeouts until the contract is written down, because without it those controls just repeat work and still can't tell you which effects already happened.
It's worth assuming the process can stop right after any durable effect. It might die between getting an API response and updating the local status, between two database commits, or after moving a file but before writing the run summary. Recovery has to treat those gaps as normal states the pipeline can end up in.