Article chapter 06 of 08
Quarantine failed records without blocking the run
One malformed record should not force a large import to restart. Skipping it without a record leaves the run impossible to reconcile. The pipeline needs to distinguish deterministic data rejection from an operational failure that may succeed later.
Validation errors include missing required fields, invalid references, unsupported values and rule conflicts. Give these stable error codes and retain the source locator. The message shown to an operator can explain the field and expected condition without exposing secrets in a general log.
Retryable faults include connection loss, rate limiting, temporary dependency failure and a worker lease expiring. Retry policy should be tied to the operation. A short database interruption may permit automatic backoff. An external action with an uncertain outcome needs reconciliation before another attempt.
Set an attempt limit, but do not make the limit the only quarantine rule. Repeating a schema mismatch several times wastes time because waiting will not change the input. Conversely, a known service outage may justify holding items in retryable until the dependency is restored without consuming attempts.
A quarantine queue should support a small number of explicit decisions: correct the source and create a new version, change a mapping under a new configuration version, mark the item intentionally skipped, or release it for another attempt. Keep the original attempt history. Editing a failed status row until it looks clean damages the record needed to understand the run.
Group repeated errors for the operator. Thousands of rows failing the same validation should appear as one dominant fault with affected counts and samples selected under privacy rules. The operator can inspect the pattern first, then open individual records for repair.