Reliable webhooks: delivery, retries, ordering and signatures

Webhooks look simple until a partner misses an event. How to design webhook delivery with outbox patterns, retries with backoff, signatures, ordering, replay and monitoring that partners can trust.

Abstract diagram of events travelling from one source to many endpoints with retry loops

A webhook is an HTTP request your system sends to a partner’s endpoint when something happens — an order is confirmed, a schedule changes, a refund completes. Sending one is trivial. Making sure every partner receives every relevant event, exactly once in effect, even when their server is down, is an engineering problem worth solving properly.

Record events before sending them

The most common failure is losing an event between changing data and sending the webhook: the database transaction commits, then the process crashes before the HTTP call. The transactional outbox pattern prevents it:

  1. In the same transaction that changes the order, insert an event row into an outbox table.
  2. A separate worker reads unsent events and delivers them.
  3. Delivery attempts and results are recorded per endpoint.

Events are never lost, because they are committed with the change that caused them.

Retry with backoff

Partners’ endpoints fail: deployments, timeouts, maintenance. Treat any non-2xx response or timeout as a failure and retry with exponential backoff and jitter over a long window — hours to days. Use short request timeouts so slow receivers do not block your workers, and deliver to each endpoint independently so one failing partner never delays others.

After the retry window, mark the event as failed, notify the partner, and keep it available for replay.

Make every event uniquely identifiable

Each event needs a unique ID that stays the same across retries. Receivers store processed IDs and ignore duplicates. Because retries mean at-least-once delivery, this is the receiver-side half of idempotency.

Sign payloads

Receivers must be able to verify that a request came from you and was not altered. A widely used approach:

  • compute an HMAC (for example SHA-256) over the event ID, a timestamp and the raw body,
  • send it in a header along with the timestamp,
  • receivers recompute it with their secret, compare in constant time, and reject requests with old timestamps.

The Standard Webhooks specification documents this pattern and has libraries in several languages. Support secret rotation by allowing two active secrets during a transition.

Ordering: design for out-of-order arrival

Retries break ordering. An “order cancelled” event can arrive before a delayed “order changed”. Rather than promising strict order, include a per-resource sequence number or updated_at, and encourage receivers to ignore events older than the state they already hold, or to fetch the current resource from your API.

Thin or fat payloads?

  • Fat events include the full resource. Convenient, but large and possibly stale on arrival.
  • Thin events include the type and resource ID; receivers fetch details. Always current, but more API calls.

A middle ground — key fields plus a link to the resource — works well for most B2B integrations.

Give partners control and visibility

  • Endpoint management with event-type subscriptions.
  • A delivery log showing attempts, status codes and response times.
  • Manual and bulk replay for a time range.
  • Test events from the dashboard or API.

These features cut support tickets dramatically, as covered in API design for B2B partners.

Monitor delivery health

Track per endpoint: success rate, latency, backlog size and age of the oldest undelivered event. Alert on backlog growth, and automatically pause endpoints that fail for long periods, with a notification to the partner.

The takeaway

Reliable webhooks rest on a few patterns: an outbox so nothing is lost, retries with backoff, unique IDs for deduplication, signatures for trust, and tooling for replay. Together they let partners build on your events with confidence.

Frequently asked questions

What makes webhooks reliable?

Recording events durably before sending, retrying failed deliveries with exponential backoff, signing payloads so receivers can verify them, giving every event a unique ID for deduplication, and offering a way to replay or fetch missed events.

How should webhook payloads be signed?

A common approach is an HMAC signature over the event ID, a timestamp and the raw payload, using a secret shared with each receiver. The receiver recomputes the signature, compares it in constant time and rejects old timestamps to prevent replay.

Should webhooks guarantee event ordering?

Strict ordering across retries is difficult to guarantee. A more robust design includes a sequence number or timestamp per resource, and lets receivers fetch the current state of the resource when events arrive out of order.