27 February 2025 / Data systems / 8 chapters

Running an item again without doing it twice

From Designing resumable large-data pipelines

Resuming only works if your effects are idempotent. The pipeline should be able to hit the same item a second time and either confirm the result that's already there or bring it to the same intended state, without creating another copy.

Database constraints protect you better than checks in application code. A SELECT followed by an INSERT can race with another worker, while a unique constraint on the source identity or business key closes that gap. Only use an upsert once you've decided what an existing row means. If you update every column on conflict, you can overwrite a correction a person made with older import data.

Before you write any retry logic, go through each kind of change the pipeline makes:

  • A create needs a stable uniqueness rule and a way to get back the identifier it created.
  • An update needs a version or comparison rule so stale source data can't quietly replace newer state.
  • An increment needs an event identity, because applying the same increment twice changes the answer.
  • A notification needs a delivery key or an outbox record.
  • A file move needs checks on both the source and destination that can tell complete, absent and ambiguous apart.

An outbox helps when a committed database change has to trigger an external action later. The same transaction writes the business change and an outbox event, and a separate dispatcher sends the event using its stable ID and records the delivery. That closes the gap where the database commits and then the process dies before the message goes out.

Idempotency doesn't mean quietly swallowing every duplicate, though. If an item turns up again with the same identity but a different input checksum, that might be a changed source, a bad identity rule or a correction from upstream. I'd stop and classify it. Treating it as a success would hide a difference that matters.

All articles