Idempotency Keys: The Cheapest Insurance in Distributed Systems

A request times out. The client has sent a payment instruction, waited its two seconds, and heard nothing back. There are exactly two possibilities: the charge never happened, or it happened and the reply was lost somewhere between the payment service and the client. From where the client sits, those are the same event. Nothing it can observe separates them.
So it has two options, and without help both are wrong. Retry, and it might charge the customer twice. Don't, and it might have dropped the order on the floor. Most code picks the first by default: the HTTP client retries on a timeout, the SDK on a connection reset, the load balancer on a 502. Most teams find out which option they picked from a customer email.
An idempotency key removes the dilemma rather than choosing a side. The client attaches a unique key to the operation, the server remembers what it did with that key, and a retry carrying the same key gets the original answer instead of a second execution. Retry as often as you like; the work happens once. It costs one table, one unique index, and one lookup per mutating request, which makes it the cheapest insurance in distributed systems by some distance. It is also easy to get subtly wrong, which is what the rest of this post is about.
sequenceDiagram
participant C as Client
participant A as API
participant D as Postgres
participant P as Payment provider
C->>A: POST /charges, key K
A->>D: claim K
A->>P: charge, key K-charge
P-->>A: charged
A->>D: store response, mark K completed
A--xC: 201 Created, lost in transit
C->>A: retry POST /charges, key K
A->>D: claim K
D-->>A: already completed
A-->>C: 201 Created, replayed
Duplicates are part of the contract
Last month I argued that RabbitMQ's at-least-once delivery is a contract rather than a defect, and that making redelivery safe is your job. That generalises well beyond brokers. Anywhere a sender has to decide whether to try again without knowing how the last attempt went, you have at-least-once semantics, whether or not anyone wrote them down.
The sources pile up quickly. Client retries after a timeout, where the request succeeded and the response never made it home. Broker redelivery, where the consumer finished the work and died before its ack landed. Publisher retries, where the broker stored the message and the confirm went missing. Proxies and load balancers that retry on your behalf, not always with your knowledge. Mobile clients replaying a queue of requests as the train leaves the tunnel. And the oldest of all: a user who clicks "Pay" twice because nothing visible happened the first time.
None of these is rare at scale, and several are correct behaviour. The retry-budget piece in July was about stopping retries from amplifying load; this one is about making them safe to send at all. A budget only makes sense for retries that are harmless. Otherwise every retry is a small wager with someone else's money, and the budget just sets the stake.
Make it naturally idempotent first
Before reaching for keys, check whether the operation can be made idempotent by its shape. RFC 9110 already classes PUT and DELETE as idempotent: setting a resource to a given state twice leaves it in that state, and deleting something twice leaves it deleted. The same trick works below HTTP. "Set the shipping address to X" is idempotent; "append X to the address history" isn't. An upsert keyed on a business identifier is idempotent; a bare INSERT isn't. "Mark invoice 1042 as paid" survives a replay. "Add £40 to the balance" does not.
Conditional writes go further. A compare-and-set on a version number, or an If-Match on an ETag, turns "apply this change" into "apply this change if nothing has moved since I looked", and a replay of a change that already landed fails the precondition instead of applying twice.
Keys are for what's left: operations that are irreducibly "do this once". Charge a card. Send a message to a human. Provision a machine. Create an order whose identifier the server assigns. These are where the money is, which is rather the point.
The key belongs to the intent, not the attempt
The most common bug in idempotency-key implementations is generating the key in the wrong place. The key identifies the intent — this customer, paying for this basket, now — and it has to exist before the first attempt. Mint it inside the retry loop and every attempt carries a fresh key, the server dutifully treats each one as new, and you have built an elaborate mechanism for charging people twice with a clean audit trail.
from uuid import uuid4
# Wrong: every attempt is a new operation as far as the server can tell.
for attempt in range(3):
resp = await client.post(
"/charges", json=body, headers={"Idempotency-Key": str(uuid4())}
)
if resp.status_code < 500:
break
# Right: one intent, one key, however many attempts it takes.
key = str(uuid4())
for attempt in range(3):
resp = await client.post(
"/charges", json=body, headers={"Idempotency-Key": key}
)
if resp.status_code < 500:
break
Where the intent outlives the process, the key has to outlive it too. A random UUID held in memory protects the retries of one call. It does nothing for a batch job that crashes and starts again from the top, because the restarted job mints new keys for the same work. There, derive the key from the business identity (f"{order_id}:capture") so that any process, restarted or not, produces the same key for the same piece of work. Stripe and the IETF's Idempotency-Key header draft (expired now, but still sensible) both suggest random UUIDs, which is right for an interactive client and wrong for a job that has to survive being run twice. Keep personal data out of derived keys, though. A key is a string, and strings end up in logs.
Two more properties matter, and both are easy to skip.
Scope the key. Keys should be unique per tenant, not globally, and every lookup must include the tenant. Otherwise two customers who happen to produce the same key collide, and in the worse case one of them is handed the other's stored response. That isn't a duplicate charge. It's a data leak with a cache in front of it.
Fingerprint the request. Store a hash of the request body with the key, and refuse a retry whose body doesn't match. A key reused with a different payload is a client bug, and quietly replaying the old response hides it. Stripe errors on mismatched parameters, AWS returns a parameter-mismatch validation error, and the IETF draft suggests a 422. Hash a canonical form, with sorted keys and fixed separators, or you'll reject honest retries from a client whose JSON serialiser has opinions about ordering.
Claiming the key
The server side is a small state machine. A key is absent, in flight, or completed, and each state, with a matching or mismatched payload, has exactly one right answer.
| Key state when the request arrives | Same payload | Different payload |
|---|---|---|
| Absent | Claim it and execute | — |
| In flight, lease still valid | 409 Conflict; retry shortly |
422 |
| In flight, lease expired | Take over, with a new owner token | 422 |
| Completed | Replay the stored response | 422 |
The claim has to be atomic, because the interesting duplicates arrive concurrently: two retries racing each other, or a retry landing while the original is still running. In Postgres that is an INSERT … ON CONFLICT, with the lease takeover folded into the same statement.
CREATE TABLE idempotency_keys (
tenant_id text NOT NULL,
key text NOT NULL,
request_hash bytea NOT NULL,
owner uuid NOT NULL,
status text NOT NULL DEFAULT 'in_flight',
locked_until timestamptz,
response_code integer,
response_body jsonb,
created_at timestamptz NOT NULL DEFAULT now(),
PRIMARY KEY (tenant_id, key)
);
-- CLAIM_SQL: returns an owner token only if this request now owns the key.
-- A takeover mints a new token, which fences out the attempt it replaced.
INSERT INTO idempotency_keys (tenant_id, key, request_hash, owner, locked_until)
VALUES ($1, $2, $3, gen_random_uuid(), now() + interval '30 seconds')
ON CONFLICT (tenant_id, key) DO UPDATE
SET locked_until = EXCLUDED.locked_until,
owner = EXCLUDED.owner
WHERE idempotency_keys.status = 'in_flight'
AND idempotency_keys.locked_until < now()
AND idempotency_keys.request_hash = EXCLUDED.request_hash
RETURNING owner;
The application wraps it:
import hashlib
import json
from collections.abc import Awaitable, Callable
from typing import Any
import asyncpg
Handler = Callable[
[asyncpg.Connection, dict[str, Any]], Awaitable[tuple[int, dict[str, Any]]]
]
class LostLease(Exception):
"""The lease expired and another attempt claimed the key."""
async def run_once(
conn: asyncpg.Connection,
tenant_id: str,
key: str,
payload: dict[str, Any],
handler: Handler,
) -> tuple[int, dict[str, Any]]:
canonical = json.dumps(payload, sort_keys=True, separators=(",", ":"))
request_hash = hashlib.sha256(canonical.encode()).digest()
owner = await conn.fetchval(CLAIM_SQL, tenant_id, key, request_hash)
if owner is None:
row = await conn.fetchrow(
"SELECT status, request_hash, response_code, response_body"
" FROM idempotency_keys WHERE tenant_id = $1 AND key = $2",
tenant_id, key,
)
if row["request_hash"] != request_hash:
return 422, {"error": "idempotency key reused with a different request"}
if row["status"] == "completed":
return row["response_code"], json.loads(row["response_body"])
return 409, {"error": "a request with this key is still in progress"}
async with conn.transaction():
# Local writes only. Calls to other systems are the next section.
code, body = await handler(conn, payload)
completed = await conn.execute(
"UPDATE idempotency_keys"
" SET status = 'completed', locked_until = NULL,"
" response_code = $3, response_body = $4::jsonb"
" WHERE tenant_id = $1 AND key = $2 AND owner = $5",
tenant_id, key, code, json.dumps(body), owner,
)
if completed == "UPDATE 0":
# Someone took the key over while this handler was still working.
# Rolling back discards the local writes it just made.
raise LostLease(key)
return code, body
Two choices in there are deliberate. The claim commits on its own, before the business transaction opens, so that a concurrent duplicate can see it. You can put everything in one transaction instead, and Postgres will make the second INSERT wait for the first to finish, which serialises duplicates for free. That works right up until the handler calls a payment provider and you're holding a transaction open across someone else's p99.
The business writes, on the other hand, commit in the same transaction as the completed marker. That is the strongest guarantee on offer: the key record and the effect it describes either both exist or neither does, so they can never disagree about what happened. Any store that can't share a transaction with your business data gives that up, and it's worth knowing what you get in exchange.
If the handler throws, the transaction rolls back and the claim stays in flight until its lease expires, after which a retry can take over. That is the right default when you don't know how far the work got. When you do know, because the request failed validation before touching anything, delete the claim so the client can retry straight away. Stripe draws the same line. It stores no result when validation fails or a concurrent request conflicts, because nothing has started executing, but it does store and replay a 500.
The lease is the sharp edge of that arrangement. A takeover assumes the previous attempt has stopped, and an expired timestamp is no evidence of that: a handler that outran its lease is still running. Without a fence you get two executions of one intent, two stored responses, and last writer wins — which is the failure the keys were bought to prevent. Hence the owner token. A takeover mints a new one, so the slow original returns to find it no longer owns the key, its completion matches no row, and the transaction it had open rolls back and takes its local writes with it. Size the lease to outlive the handler's own timeout so this stays the rare path, and renew it mid-flight if the work is genuinely long. What no fence can undo is a call already made to another system, which is the next section.
What to keep, and for how long
What you store is the response, verbatim: status code and body. A retry should be indistinguishable from the original reply, including the identifier the server minted on the first attempt, because that is precisely the thing the client lost.
How long you keep it is where most implementations spring a quiet leak. The TTL has to exceed the longest gap between an attempt and its last possible retry, and that gap is rarely the one in the client's backoff config. Stripe may prune keys once they're 24 hours old, and says plainly that a key reused after pruning is treated as a new request. SQS FIFO queues deduplicate over a five-minute window, which is generous for a client retry and useless for anything slower. The classic hole is the dead-letter queue: messages park on Tuesday, someone fixes the bug and replays them on Friday, and every key that expired in between turns the replay into a second execution.
For consumers, then, a TTL'd key store is the weaker tool. Where the message carries a business identity, put a unique constraint on the effect itself, such as one ledger row per (order_id, event_type), and let the database refuse the duplicate however late it turns up. Keys expire. Business state doesn't.
The choice of store shapes the rest:
| Store | Atomic claim | Shares a transaction with the business write | What bites |
|---|---|---|---|
| Postgres, same database | INSERT … ON CONFLICT |
Yes | Table growth; needs a sweeper |
| Redis | SET key value NX PX ttl |
No | Eviction under memory pressure silently forgets keys |
| DynamoDB | Conditional put on attribute_not_exists |
Only via TransactWriteItems |
TTL deletion is a lazy background sweep; check the timestamp |
| Process memory | A dict | No | Forgets everything on restart; not idempotency at all |
Redis earns a specific warning. It's fast, it's obvious, and it's usually already there — as a cache, configured with allkeys-lru. Put idempotency records on that instance and, the first time memory gets tight, Redis will evict them, correctly, according to the policy you gave it. Nothing errors. The duplicates simply start getting through. Keep them on an instance set to noeviction, or accept that your insurance lapses whenever you're busy.
Partial failure is where it gets expensive
The handler above sticks to local writes for a reason. The moment it touches something outside your database, whether a payment provider, an email service, or another team's API, the transaction no longer covers the work, and a crash in the wrong place leaves the key in flight with an effect that has already escaped.
Take a checkout that writes the order, charges the card, and sends a receipt. The process dies after the charge succeeds and before anything is recorded. The lease expires, the client retries, the retry takes over the key, and it runs the charge again, because nothing in your database says it happened.
There are two fixes, and serious systems use both.
The first is to propagate idempotency downstream. When your handler calls the payment provider, it sends a key of its own derived from yours, f"{key}:charge", so the replayed charge is deduplicated by the provider and returns the original result. Idempotency only composes if every hop passes a key along. The one hop that doesn't is where your duplicates come from.
The second is to make the handler resumable. Brandur Leach's write-up of Stripe-style keys in Postgres, still the best treatment I know of, splits a handler into atomic phases — runs of local writes, each committed in its own transaction — separated by the calls to other systems, which he calls foreign state mutations. After each phase the key record stores a recovery point, and a retry that takes over the key resumes from the last one rather than starting at the top. The foreign calls still need keys of their own; recovery points just stop you making those calls more often than necessary.
Some side effects can't take a key: most email, most SMS, anything that ends at a human. Those get a different treatment. Write the intent to an outbox table in the same transaction as the business change, and let a relay deliver it at least once. You won't get exactly one receipt in every failure case; you'll get the occasional duplicate. That is the right trade, provided you've decided per side effect what a duplicate costs. A second receipt is mildly embarrassing. A second charge is a refund, a support ticket, and a customer who now reads their statements carefully.
One thing keys don't solve is ordering. A key stops a request applying twice; it doesn't stop a delayed retry of an old request landing after a newer one. If that matters, it's a job for version numbers and conditional writes, not a bigger dedupe table.
Exactly-once is a property of the effect
Tyler Treat's "You Cannot Have Exactly-Once Delivery" makes the case more thoroughly than there's room for here. Over an unreliable network, a sender can't distinguish a lost message from a lost acknowledgement, so delivery is at-most-once or at-least-once and nothing else is on the menu. What systems advertise as exactly-once is at-least-once delivery plus an effect that absorbs duplicates. Kafka's exactly-once semantics are genuine and genuinely bounded: they cover reading from and writing to Kafka's own log. The moment your consumer calls anything else, you're back to keys and constraints.
That isn't a disappointment. It's the design. Once you stop trying to deliver things once and start making the effect idempotent, retries stop being a correctness question and become a policy one — how many, how fast, within what budget — which is a far better argument to be having in an incident review.
I run usage accounting for an LLM gateway on at-least-once workers, where a duplicate isn't a log line; it's a line on somebody's invoice. That concentrates the mind. One table, one unique index, one lookup per request: that's the whole premium, and it's paid long before anyone needs to make a claim.
Further reading
- Brandur Leach, Implementing Stripe-like Idempotency Keys in Postgres
- Malcolm Featonby, Making retries safe with idempotent APIs (Amazon Builders' Library)
- Stripe, Idempotent requests
- IETF httpapi working group, The Idempotency-Key HTTP Header Field (expired draft)
- Tyler Treat, You Cannot Have Exactly-Once Delivery