17 September 2026 / Integrations

Checking an integration after the job says it finished

A long agent run can complete most of its work and still miss an important requirement. I want the same distinction in production integrations: a completed run and a verified result are separate facts.

An empty industrial service corridor with cables and machinery running along one wall.
Generated image: Villar David editorial

Android Bench 2.0, published this week, adds long-horizon coding tasks and reports completion separately from pass rate. An agent can complete most of a multi-day task and still miss an important requirement. I want the same distinction in production integrations: a completed run and a verified result are separate facts. Consider a dashboard that receives account updates from another application. The scheduled job runs, the dashboard shows a recent update time, and someone opens an account that still has yesterday's details. That is a hypothetical example, but it gives us a useful place to start. Before rerunning the job, I would check what its success status actually means. It might mean the source request returned without an error. It might mean every fetched record was placed on a queue. Neither tells us whether the destination accepted the changes. If the dashboard reads from a separate reporting table, there is another step before the user sees anything. Each part can finish successfully while a later part is waiting or has rejected a record.

The status needs to describe the stage it belongs to. "Source checked" is useful. So is "updates waiting to be applied". A single "last synced" timestamp is hard to interpret when it could refer to either of those events. I would keep the detailed timings in the run record and show the user the one that answers their question about the data on screen. There is a similar issue with incoming webhooks. Stripe's webhook documentation asks endpoints to return a successful response quickly, before complex processing. It also documents duplicate deliveries and says event delivery order is not guaranteed. An acknowledgement therefore needs a precise meaning in your own application. If it means the event has been durably accepted for processing, the remaining work needs its own status and recovery path.

For the dashboard example, I would start with a small comparison against the source. Use stable record identifiers and inspect the fields the integration owns. Names are poor matching keys, and a total record count can agree while individual records differ. A missing account and a duplicate account can cancel each other out in the count. Keep the comparison narrow enough to understand. There may be fields that staff deliberately maintain in the destination, or a transformation that changes how values are stored. Document those rules before calling every difference a defect. Otherwise a repair job can overwrite a legitimate local change while trying to make two systems identical. I would want the result to show:

  • source records that should exist in the destination but are missing;
  • records whose mapped fields differ;
  • records waiting for processing, including how long they have waited;
  • rejected records and the reason they could not be applied;
  • records that could not be compared because a lookup failed.

That last group matters. If the comparison cannot read the destination, the result is incomplete. It should say so. Quietly treating an unavailable record as absent can turn a temporary read problem into a duplicate creation attempt. Freshness needs a little care as well. The time of the most recent changed record tells you when that record changed. It doesn't tell you whether the integration has checked for new work since then. A quiet system may have old records and a healthy integration. A busy system may contain one very recent record while older updates remain stuck in the queue. So I would keep the latest completed source check alongside the age of the oldest unresolved item. Where the source supports a cursor or sequence, record how far processing has reached and whether there are gaps behind that position. Moving a cursor past a rejected record is a design choice that needs an explicit recovery mechanism. It should not make the rejected work disappear from the operator's view. Time windows also need to match. Comparing an export taken before a change with a live destination read taken afterwards will produce differences that may be perfectly valid. Use a source snapshot where one is available, or record the comparison window and recheck differences before acting. Be explicit about timezone conversions, especially when a date filter selects the records to inspect.

Once there is a clear list of differences, the repair can be scoped to those records. Some may need replaying. Others may need corrected source data or a decision about which system owns the field. I would keep uncertain cases out of an automatic repair batch and show them separately for review. After the repair, run the comparison again. Record the remaining differences, including anything deliberately excluded. A repair job finishing is another process result; the comparison supplies the evidence about the records. For a high-volume integration, a complete scheduled reconciliation may need to run in batches, with smaller checks during the day. A sample can help detect problems, but it only supports a claim about the records actually examined. I would put one manageable check on the next maintenance task: choose an integration that currently reports a last-run time, trace a record through to the screen that consumes it, and write down where each status is set. Then add a read-only comparison for the records that should have arrived. Leave its unresolved results visible until someone has checked them.

Continue the thinking.

Comments are public and hosted in an open-source GitHub Discussions repository.

Loading comments connects your browser to GitHub. A GitHub account is required to post.

All blogs