30 January 2025 / Data operations

Long-running data jobs should be designed to stop

Large imports rarely fail at a convenient boundary. They need deliberate stopping points, a record of what completed and a restart path that does not duplicate accepted work.

Gold contact pads and fine traces on a dark printed circuit board.
Photo: Vishnu Mohanan (opens in a new tab)

Large imports rarely fail at a convenient boundary. They need deliberate stopping points, a record of what completed and a restart path that does not duplicate accepted work. If a data job can only run uninterrupted, every dependency must remain healthy for its full duration. The database connection cannot drop. Credentials cannot expire. The process cannot be redeployed. Nobody can discover a bad source file halfway through. For a job that may run for hours, those are poor assumptions. Assume the process will be interrupted and decide how it stops and resumes. The first decision is the smallest useful unit the job can accept or reject. It might be one source file, one account, a page of records or a bounded ID range. "We processed 43 per cent" is not enough. The job needs to say which units reached a committed state. A useful unit has a stable identity. If a page number changes when records are added, it is a poor restart key. A source record ID, file checksum or explicit range is more dependable. The job can record that identity before work begins and mark it complete only after every related write succeeds.

Database transactions help, but I would not stretch one transaction over the whole import. Long transactions retain locks and make recovery expensive. A transaction around one bounded unit gives the database a realistic rollback boundary and gives the operator a comprehensible checkpoint. The unit should also be large enough to avoid spending most of the runtime on coordination. Start with a conservative batch size, measure it, then adjust. The important bit is that changing the batch size does not change the meaning of completion. A restart will revisit something. The system should decide what that means before it happens. For a simple insert, a unique source key can prevent a duplicate. For an update, the job may use an upsert based on that source key. For a multi-table import, it may need a small state machine: received, validated, written and reconciled. Whichever method is used, processing the same accepted unit twice must not create a second logical result.

Blindly skipping every existing row is risky. An earlier attempt may have written only part of the record, or the source may contain a legitimate correction. I prefer to store enough information to distinguish these cases. A source version, content hash or import run identifier can show whether the incoming item matches what was accepted earlier. Side effects outside the database need the same treatment. If the job sends a message, creates a file or calls another system, a database rollback cannot undo that action. Give the side effect its own idempotency key where the receiving system supports one. Otherwise, put it behind an outbox or a separate dispatch step that records delivery attempts. An in-memory counter is useful for a progress bar and useless after a crash. Checkpoint state belongs in durable storage that survives the worker. For each unit, I want to know:

  • the stable unit identifier;
  • the import run and source version;
  • when processing started and finished;
  • the final state, such as accepted, rejected or awaiting retry;
  • a short error code with a link to fuller diagnostic detail;
  • the code or schema version that interpreted the input.

The restart query should read this record directly because it controls what work resumes. If the process says "resume from batch 380" while the durable checkpoint says batch 379 never committed, resume from the durable checkpoint. Keep job state separate from business data where practical. That makes it easier to inspect an interrupted run without reverse-engineering half-written domain records. It also gives operators somewhere to record an intentional pause rather than making a stopped process look identical to a failure. Retries are useful for transient failures: a connection timeout, rate limit or short service outage. They are harmful when the input itself is invalid. Retrying a malformed date ten times only delays the rest of the file. Classify errors close to where they occur. A validation failure can move to a rejection queue with the source identifier and reason. A transient dependency failure can use bounded retries with increasing delay. An unknown exception should stop the affected unit and raise an alert rather than being silently converted into a generic rejection. This distinction matters during restart. The operator can resume transient work without feeding known bad records through the same path again. Rejected records can be corrected and submitted as a new source version, preserving the earlier decision rather than overwriting history. A retry limit also prevents one poisoned unit from holding a worker forever. Once the limit is reached, the job should leave an explicit failed state. Someone can then decide whether to correct, skip or investigate it.

Sometimes nothing has failed. A deployment is due, a source was wrong or the database needs maintenance. Sending a hard kill may leave the current unit in an uncertain state. A cooperative stop is cleaner. The worker receives a stop request, finishes or rolls back its current transaction, records the checkpoint and exits. New units are not claimed after the request. The operator gets a final status that says where the run stopped and whether any work remains in progress. Put time limits around shutdown as well. A worker stuck on an unresponsive call cannot promise to stop cleanly. Network calls need timeouts, database queries need sensible limits and the supervisor needs a final hard-stop option. The difference is that the hard stop is a fallback with a known recovery procedure. Interrupt the job on purpose before trusting its restart path. Run a representative sample and stop it during validation, during a database write and after the write but before the completion marker. Restart each case. Check the business records, checkpoint table, rejected items and any external side effects. Then run the reconciliation query that compares accepted source units with committed destination records. One final production check is to choose a completed unit and submit it again. The data should remain logically unchanged, and the job record should explain what it did. If that test is ambiguous, the large import is not ready to be left alone.

Continue the thinking.

Comments are public and hosted in an open-source GitHub Discussions repository.

Loading comments connects your browser to GitHub. A GitHub account is required to post.

All blogs