Article chapter 04 of 08
Put checkpoints after durable effects
Place a checkpoint only after everything before it is durable and safe to treat as complete. Updating a counter in memory or writing a log message before the destination commit creates a progress record without the corresponding result.
For database-only work, keep the destination write and item acceptance in the same transaction where possible. The sequence is straightforward: validate the item, perform the destination mutation, record destination references, mark the item accepted, then commit. An interruption before commit leaves no accepted item and no destination change. An interruption after commit leaves both.
External systems break that neat boundary. An API call cannot usually participate in the local database transaction. Use an idempotency key if the destination supports one, and record enough information to reconcile an uncertain response. The flow may be:
- Create an attempt with a stable idempotency key.
- Send the external request with that key.
- Store the returned external identifier and response classification.
- Apply any local write.
- Mark the item accepted.
If the worker stops after step two, the next attempt should query or repeat the call using the same key rather than generate a fresh action. Where the destination has no idempotency mechanism or lookup, route uncertain items to reconciliation. Blind retry is unsafe when the first request may have succeeded.
Checkpoint frequency is an operating choice. Per-item commits make recovery precise but may cost too much for high-volume database work. Batch commits improve throughput, though interruption replays the uncommitted batch. Choose a batch size based on transaction duration, lock pressure, memory use and acceptable replay, then record the committed boundary. Do not base it only on the number that looked fast in a local test.