12 March 2024 / Payment workflows

The workflow that starts when a payment clears

Recording a successful payment is only one part of the job. The surrounding workflow has to update the right account, send the right message, avoid duplicate handling and leave enough state for support staff to work out what happened when one step fails.

Intersecting steel beams and bracing viewed from beneath a bridge.
Photo: Sebastian Schuster (opens in a new tab)

A payment provider can notify an application that a transaction succeeded. The handler should accept that event once, record it and start the downstream work from local state. A repeated notification should receive a successful acknowledgement without repeating completed actions. Webhooks can be retried, delivered out of order or received after a user has closed the checkout page. The browser redirect cannot be the authoritative signal either. A customer may pay successfully and never return to the application, or refresh a completion page several times. The first job is to validate the incoming event. Check its signature using the provider's documented method, parse it only after verification, and confirm that the event type is one the application expects. Store the provider's event identifier and the local time it was received. If the same identifier arrives again, the handler should recognise it and return a successful acknowledgement without redoing the work. After acceptance, the workflow can use internal state and its own retry rules. The application needs a reliable link between the external transaction and the account, order or subscription it affects. Email address is usually a weak key because it can change, vary in case, or belong to someone buying on another person's behalf.

Create the local order or payment attempt before sending the user to checkout. Pass an opaque internal reference through the provider's supported metadata or reference field. When the notification returns, resolve that reference and verify the facts that should agree, such as currency, amount and provider account. Do not let metadata silently override authoritative local data. Its purpose is correlation. The application should already know what the person attempted to buy and the expected amount. A mismatch belongs in an exception state for review rather than being forced through the normal success path. Keep both identifiers on the payment record: the local identifier used by the application and the provider's transaction identifier used for reconciliation. Support staff will need one or the other depending on where a question begins. Once the payment is accepted, several things may follow. Access may be granted, an order may be marked paid, a receipt may be requested, and a message may be queued. Any of those steps can fail after an earlier one succeeds. A single database transaction can protect changes inside one database, but it cannot usually include an email service, accounting system and payment provider. The workflow therefore needs idempotency at each boundary. Granting access twice should produce the same final entitlement. Retrying a message job should not send a second receipt without an explicit reason. Creating an external record should use a stable reference or store the returned identifier before another attempt.

An outbox pattern is useful for work that leaves the main database. In the same transaction that marks the payment accepted, write durable jobs describing the required downstream actions. A worker can process those jobs, record attempts and retry failures. This avoids the gap where the database commits but the process crashes before a message reaches the queue. The job payload should identify the payment and intended action, not copy every mutable account field. The worker can read current authorised data when it runs. Store “paid” and “confirmation email sent” as separate states. Combining them makes partial failure hard to describe. Use explicit fields or related event records for the parts operators need to inspect. A payment workflow may need to distinguish:

  • event received and verified;
  • payment matched to an internal record;
  • payment accepted or held for review;
  • account or entitlement updated;
  • customer notification queued, sent or failed;
  • external accounting or fulfilment update pending or complete.

These states do not need to be exposed in full to the customer. Internally, they prevent a failed email from making a successful payment look unpaid, and they prevent a paid flag from implying that every consequence completed. Transitions should record a timestamp and reason. If an operator changes a held payment after checking it, keep that action distinct from the automated event. This makes later reconciliation possible without reading application logs line by line. The success message should be generated from the accepted internal record, not directly from an unverified webhook body. This keeps the amount, product description, account name and next steps consistent with what the application has recorded. Queue the message after the relevant account or access update commits. If the customer is told that access is ready while the entitlement job is still pending, the message creates a support problem. Where the update can take time, say that plainly and provide a status the application can support. Use a stable message key tied to the payment and message type. The sending service or local job table can check that key before another send. A support operator may still need a deliberate “resend” action, but that action should create a new recorded attempt rather than bypassing duplicate protection.

Messages should avoid exposing internal error detail. A failed downstream integration belongs in an operator queue. The customer-facing status should explain what they can do now and how to get help, using the facts the system can confirm. When one step fails, support needs to answer a small set of questions quickly. Did the provider accept the payment? Which local record did it match? Which downstream actions completed? What is safe to retry? A useful internal view searches by local payment ID, provider transaction ID and account reference. It shows the amount and currency, verified event history, current business state, downstream jobs, latest error and available operator actions. Raw payloads may help technical investigation, but they should be access-controlled and should not be the main support interface. Retries need narrow controls. “Retry notification” is safer than “run payment workflow again” because it states which incomplete action will occur. The handler should still enforce idempotency rather than trusting the button. Reconciliation also needs a periodic check independent of webhooks. Compare provider transactions with local payment records and list unmatched, mismatched or incomplete items. During implementation, replay the same verified event twice and interrupt the workflow after each downstream step. After recovery, check the account, message count and payment state separately.

Continue the thinking.

Comments are public and hosted in an open-source GitHub Discussions repository.

Loading comments connects your browser to GitHub. A GitHub account is required to post.

All blogs