Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Retries Should Have a Budget—Not Just a Backoff

A telegraph-window teller spending from a small dwindling stack of brass retry-tokens as a queue waits and a slow pendulum marks the wait — a finite budget, not just backoff — in 1960s gouache.

Every incident review I've sat through where retries made things worse ends the same way: someone proposes adding exponential backoff and jitter, everyone nods, and we move on. Job done. Except it isn't.

Backoff doesn't stop retry amplification. It just spreads it over a slightly longer window. If your downstream is already on fire, taking 800ms instead of 200ms to pile on doesn't save it — it just changes the shape of the graph.

The maths is unforgiving. With K retries per call across N service hops, the deepest service sees up to K^N times the normal load. Three retries, three hops deep — that's 64x. Your database doesn't care that the requests are jittered. It cares that there are 64 of them.

Treat retries as a finite resource

The fix isn't a smarter backoff curve. It's a budget. Pick a number — something like "retries may consume no more than 10% of total request volume" — and enforce it client-side. When the budget is exhausted, you stop retrying. Full stop. The request fails fast and the error propagates honestly.

This sounds harsh until you realise the alternative: retries continue indefinitely, the downstream stays saturated, and every layer above piles on more work hoping the next attempt will land. That's not resilience. That's a self-inflicted DDoS dressed up in middleware.

gRPC has had retry budgets in the spec for years. Envoy supports them. Finagle pioneered the pattern. Yet most internal HTTP clients I encounter still ship with naive retry-on-5xx and call it resilient.

Tie the budget to your SLO

Here's the bit people skip: the budget number isn't arbitrary. If your SLO allows 0.1% errors, a 10% retry budget gives you 100x headroom for transient failures before retries themselves become the problem. If you're already burning error budget, retries should be the first thing throttled, not the last.

Retries are borrowed capacity. Backoff decides when you borrow. A budget decides whether you're allowed to borrow at all. Build the second one, and the first matters a lot less.

If your client library can't express a retry budget, that's the bug. Not the latency on your downstream.