Article chapter 04 of 08
Where checkpoints should go
I'd only put a checkpoint after everything before it is durable and safe to count as done. If you bump a counter in memory or write a log message before the destination commit, you've got a progress record for a result that doesn't exist yet.
For work that only touches the database, keep the destination write and the item acceptance in the same transaction wherever you can. The order is simple: validate the item, make the destination change, record the destination references, mark the item accepted, then commit. If it stops before the commit, you've got no accepted item and no destination change. If it stops after, you've got both.
External systems mess that up. An API call usually can't take part in your local database transaction. Use an idempotency key if the destination supports one, and record enough to sort out a response you're not sure about. The flow might look like this:
- Create an attempt with a stable idempotency key.
- Send the external request with that key.
- Store the external identifier that comes back and how you classified the response.
- Make any local write.
- Mark the item accepted.
If the worker stops after step two, the next attempt should look up or repeat the call using the same key, rather than kicking off a fresh action. If the destination has no idempotency support and no way to look things up, send those uncertain items to reconciliation instead. Retrying blindly isn't safe when the first request might have worked.
How often you checkpoint is an operating decision. Committing per item makes recovery precise, but it can cost too much for high-volume database work. Batch commits are faster, though an interruption means replaying whatever batch hadn't committed. I'd pick a batch size based on how long transactions run, lock pressure, memory use and how much replay you can live with, and then record the committed boundary. The number that happened to look quick in a local test shouldn't be the whole basis for it.