Skip to content
JuicePay
All posts

API

Designing idempotent payout APIs for retry storms

Marcus Lindqvist · · 9 min read

The failure we designed hardest against is not a network outage. It is a client that means well.

Picture the sequence: your service submits a 10,000-leg payroll batch. The request succeeds on our side in 900ms. The response is lost somewhere between our edge and your process, because a load balancer recycled a connection. Your client times out at ten seconds and retries. If that retry is not correctly keyed, you have just paid everyone twice — a $1.2M mistake that takes weeks to unwind.

Idempotency is a contract, not a header#

It is tempting to treat Idempotency-Key as a flag you set to feel safe. It is really a contract with three parties, and all three have obligations.

The server (us) must:

  • Persist the key alongside the resulting resource before responding.
  • Return the stored response verbatim on a repeat, including the original status code.
  • Detect a key reused with a different body and refuse loudly.

The client (you) must:

  • Derive the key from your own domain, not a random number.
  • Keep it stable across every retry of the same logical action.
  • Treat a 409 as a bug in your code, not a transient error.

The operation must be genuinely safe to repeat. This is the obligation people forget. If the underlying action has side effects that cannot be made idempotent, no amount of keying will save you.

The keying trap#

Here is the version of this that looks correct and is not:

ts
// Looks fine. Pays everyone twice on retry.
const key = `payroll:${run.id}`;
await Promise.all(
  recipients.map((r) =>
    juicepay.payouts.create(
      { amount: r.amount, asset: "USDC", recipient: r.walletId },
      { idempotencyKey: key },
    ),
  ),
);

The first request creates a payout. Every subsequent request with the same key returns that first payout — for a different recipient. The remaining 9,999 recipients either fail or, worse, get ignored, and your payroll run pays one person.

The key must identify the logical operation, and the logical operation here is "pay recipient X in run Y":

ts
const key = `payroll:${run.id}:recipient:${recipient.id}`;

Now each recipient has a stable, unique key. Retrying the whole Promise.all is safe: each request either creates its payout or returns the one already created, and the outcome is identical either way.

Batches change the shape of the problem#

Per-record keys work for fan-out, but a batch endpoint has a different concern: a single HTTP request creates ten thousand resources. You need batch-level idempotency, and you need leg-level recovery.

Our approach is to derive leg keys deterministically from the batch key and the row index:

text
leg_key = hmac(batch_key, row_index)

This gives you three properties that matter:

  1. A retried batch returns the existing batch, not a new one.
  2. An individual leg can be replayed safely — the derived key is identical, so a leg that already settled returns its original result.
  3. You never have to know which subset landed. Submit the same batch, get the same answer.

Property two is the one that pays for itself. When three legs out of ten thousand fail permanently — a bad address, a blocked recipient — your recovery loop is:

ts
const failed = await juicepay.payouts.listLegs(batch.id, { status: "failed" });
await juicepay.payouts.replayLegs(batch.id, {
  legs: failed.data.map((leg) => ({
    row: leg.row,
    recipient: correctedRecipientFor(leg.reference),
  })),
});

No reconciliation against a partial result set, no manual bookkeeping of which legs you already retried. The deterministic key does the work.

Why we reject rather than ignore#

When a key arrives with a body that differs from the original request, we return 409 idempotency_key_reuse. The alternative — returning the original response and ignoring the new body — is worse, because it hides a real bug. A key collision means either your key derivation is not deterministic, or two genuinely different operations are sharing a key. Both are defects, and both should surface in development rather than in production three months later.

The same logic drives our response to malformed keys. We validate that the key is a non-empty string under 255 characters, and we normalise case before hashing. That last detail came from a real incident: a retry that uppercased the key created a duplicate payout, because the hash was case-sensitive and the two keys looked identical to a human.

Interaction with webhooks#

Idempotency secures the API boundary. It does nothing for your webhook consumer, which faces the same problem from the other direction: at-least-once delivery means duplicates are normal, not exceptional.

The consumer-side pattern is a claim, not a check:

sql
INSERT INTO processed_events (id, type, received_at)
VALUES ($1, $2, now())
ON CONFLICT (id) DO NOTHING
RETURNING id;

An empty result means another delivery already claimed the event. A read-then-write check — SELECT then INSERT — races under concurrent deliveries, which is precisely when duplicates cluster. Use the insert's return value as the lock.

The retention question#

Keys are retained for 24 hours. That window is a design decision with an operational consequence: if your retry strategy can exceed 24 hours, a retry will create a new resource.

For automated retries with bounded backoff, 24 hours is generous. For a manual workflow — a finance operator clicking "resend" the next morning — it is not. The fix is not a longer window; it is to persist the resulting resource ID on your side and check it before re-requesting. The key protects the automated path, and your own record protects the human one.

What we got wrong first#

Our first implementation stored the key on the resource rather than in a dedicated table, which meant a lookup required knowing which resource type to search. That worked until a batch and a single payout shared a key, at which point the lookup returned whichever type we happened to check first.

The fix was a key table separate from every resource, indexed by key and scoped by account and endpoint. It is less elegant than embedding the key in the resource and it is the only version that survives a support engineer trying to answer "did this request actually create anything?" at 2am.

The general lesson, which applies well beyond payments: idempotency is a property of the system, and systems that try to keep their guarantees in more than one place eventually disagree with themselves.

Want to try this against a real API?

Sandbox keys are issued instantly and settle against deterministic fixtures, so nothing in this post requires production funds to reproduce.