27 February 2025 / Data systems / 8 chapters

Reprocess items without duplicating effects

From Designing resumable large-data pipelines

Resumption depends on idempotent effects. The pipeline should be able to encounter the same item again and either confirm the existing result or bring it to the same intended state without creating another one.

Database constraints provide stronger protection than application checks. A preliminary SELECT followed by INSERT can race with another worker. A unique constraint on the source identity or business key closes that gap. Use an upsert only after deciding what an existing row means. Updating every column on conflict can overwrite a newer human correction with older import data.

Classify each mutation before implementing retry:

  • A create needs a stable uniqueness rule and a way to recover the created identifier.
  • An update needs a version or comparison rule so stale source data cannot silently replace newer state.
  • An increment needs an event identity, because applying the same increment twice changes the result.
  • A notification needs a delivery key or an outbox record.
  • A file move needs source and destination checks that distinguish complete, absent and ambiguous states.

An outbox is useful when a committed database change must trigger a later external action. The same transaction writes the business change and an outbox event. A separate dispatcher sends the event using its stable ID and records delivery. This avoids the gap where the database commits but the process stops before publishing a message.

Idempotency does not mean swallowing all duplicates. A repeated item with the same identity but a different input checksum may indicate a changed source, a bad identity rule or an upstream correction. Stop and classify it. Treating it as success would hide a material difference.

All articles