Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

SLOs Without the Theatre: Error Budgets That Survive Contact With Management

A grand near-empty theatre where a single large dial on an easel is spotlit on stage, admired by a few distant figures, its cable trailing off the stage edge to an unplugged connector beside an idle machine lever — ceremony wired to nothing, in 1960s gouache.

Somewhere in your organisation there is a dashboard with a number on it that says 99.9%. It is green. Nobody looks at it. When the service last fell over, nobody consulted it before deciding what to do, and nobody consulted it afterwards to decide whether the decision had been correct. The number exists because a planning cycle two years ago had a line item called "define SLOs", and that line item got a tick.

That is the natural resting state of an SLO programme. It produces an artefact, the artefact is admired briefly in a quarterly review, and then it does precisely nothing, because at no point was it wired to a decision anyone was obliged to honour. An SLO that changes no behaviour is not an SLO. It is a graph.

The version that survives a quarter is the version that occasionally tells people something they do not want to hear, and is believed anyway. Everything below is in service of that.

What an SLI actually measures

The first place these programmes go wrong is the indicator, and they go wrong in a specific, recognisable way: the team measures what the server did rather than what the user experienced. This is the request-versus-user trap, and it is seductive because request-level metrics are sitting right there in your Prometheus instance, already labelled, already aggregated, asking to be turned into a ratio.

Consider a checkout service. The obvious SLI is the proportion of HTTP requests that return a 2xx. It is easy to compute, it produces a satisfyingly high number, and it is very nearly meaningless. A single user checking out makes a dozen requests; if eleven succeed and the twelfth — the one that actually takes the payment — returns a 500, your request-success ratio reads 91.7% for that user, which sounds like a bad five minutes. The user's experience is that checkout is down. Those are not the same fact, and the gap between them is where your reputation quietly leaks away.

The indicator has to track the thing the user came to do, end to end, as a unit. Sometimes that is a single request and the simple ratio is honest. Often it is a journey across several requests, or a queue depth, or the freshness of a value rather than the success of a call. The work is in deciding what "good" means from outside the system, then finding the closest measurable proxy for it — not in finding a measurable thing and declaring it good because it was to hand.

Vanity SLI User-centric SLI Why the difference matters
Proportion of HTTP requests returning 2xx Proportion of completed checkout journeys with no failed step A user lives in journeys; one fatal request among twelve still ruins the journey
Average request latency across all endpoints 99th-percentile latency on the three endpoints on the critical path Averages hide the tail, and the tail is where users actually suffer
Job-queue worker uptime Proportion of jobs completed within their freshness deadline A worker that is up but four hours behind is, to the user, down
API gateway availability Availability as observed from the client, including DNS, TLS, and edge The user's failures include the half you don't host

The right-hand column is harder to instrument. That is not a coincidence. The metrics that correlate with user pain tend to live at boundaries you don't fully control, which is exactly why the easy metric and the honest metric diverge.

Picking the three that matter

You do not need an SLI for everything, and the instinct to define one per endpoint is how you end up with forty indicators, no error budget anyone can reason about, and a spreadsheet that becomes its own small bureaucracy. Pick three. Occasionally four. Past that, nobody — including you — holds the whole picture in their head, and an SLO you cannot hold in your head is one you will not act on under pressure.

The three almost always map to the same shape of question. Is the thing available when a user reaches for it. Is it fast enough that the user doesn't give up. Is the result correct, or fresh, or otherwise not quietly wrong. Availability, latency, and a quality dimension — the last being the one teams skip, because it is the hardest to measure and the most embarrassing to get wrong.

Quality is where the interesting failures hide. A search service that returns results quickly and reliably, all of which are stale by six hours, scores beautifully on availability and latency and is nonetheless broken. A data pipeline that is up and responsive while silently dropping one row in a thousand is the kind of failure that doesn't page anyone for a year and then costs you an audit. If your three SLIs are all availability and latency, you have measured whether the system is responding, not whether it is right, and the difference is the entire reason anyone employs you.

Choose the three by asking which failures would make a user leave, escalate, or file a complaint they mean. Then measure those, and let the rest go uninstrumented at SLO level. They can still have alerts. They just don't get a budget.

Error budgets as a contract

Here is where the SLO stops being a measurement and starts being an agreement, and the shift is the whole point. A target of 99.9% over a rolling 28 days is not a description of how the system behaves. It is a statement that you are permitted to be unavailable for roughly 40 minutes in that window, and that within those 40 minutes nobody gets to be cross with you. Above the line you have spent the budget, and the rules change.

That inversion is what makes error budgets useful and what makes people uncomfortable. A reliability target framed as a goal invites the response "well, more reliable is always better, push it to five nines". A reliability target framed as a budget makes the cost legible: those nines are paid for in engineering time, deferred features, and on-call attrition, and a budget you have deliberately not spent is a budget you have wasted on reliability the user could not perceive. Nobody thanks you for the latency you shaved below the threshold of human noticing.

For the contract to mean anything, three things have to be true, and they are usually the three that get quietly dropped. The number has to be agreed by the people who will be bound by it, not handed down — a budget imposed by a platform team on a product team is a budget the product team will spend with enthusiasm and no remorse. The window has to be a rolling one, not calendar-aligned, because a budget that resets cleanly on the first of the month invites a particular kind of recklessness on the 28th. And there has to be a written, agreed consequence for exhausting it. Without the third, you have a target with good manners and no teeth.

Burn-rate alerting, fast and slow

Once the budget is real, the question becomes how fast you are spending it, and this is where most alerting setups are either useless or actively hostile. The naive approach pages you the moment the success ratio dips below the SLO target. That fires constantly, because a service sitting exactly at 99.9% will breach over short windows all the time through ordinary noise, and within a fortnight everyone has a mail filter routing the alert straight to the bin. The alert that cried wolf is worse than no alert, because it trains the team to ignore the channel the real one will arrive on.

Burn rate fixes this by alerting on the speed of consumption rather than the level. A burn rate of 1 means you are spending budget at exactly the pace that exhausts it precisely at the end of the window — sustainable, by definition. A burn rate of 10 means you will exhaust a 28-day budget in under three days. That is the number worth waking someone for.

You need two regimes, and they answer different questions. A fast-burn alert — a high burn rate sustained over a short window, say 14x over an hour — catches the acute incident: something broke at 02:00 and is eating budget at a rate that will blow the whole month by breakfast. That one pages. A slow-burn alert — a more modest rate, perhaps 3x sustained over six hours — catches the chronic degradation that no single incident would trip: the slow leak, the gradual latency creep, the dependency that got 2% worse after a deploy and stayed there. That one does not need to wake anyone. It needs to create a ticket and ruin someone's Tuesday in daylight, which is a far more humane way to lose an argument with entropy.

The multi-window refinement — requiring the burn rate to hold over both a long and a short window before firing — exists to kill the false positive where a 90-second blip trips the short window and resolves before you have found your laptop. Both windows must agree. It costs you a little detection latency and buys you an on-call rotation that still trusts its pager, which is a trade worth making every time.

YesYesNoNoYesYesNoNoCompute burn rateover short and longwindowsShort-window burn >=14x?Long-window burn >=14x?Page on-call now —fast burn, acuteincidentHold — likely atransient blip, keepwatchingShort-window burn >=3x?Long-window burn >=3x?Open ticket — slowburn, chronicdegradationWithin budget — noactionYesYesNoNoYesYesNoNoCompute burn rateover short and longwindowsShort-window burn >=14x?Long-window burn >=14x?Page on-call now —fast burn, acuteincidentHold — likely atransient blip, keepwatchingShort-window burn >=3x?Long-window burn >=3x?Open ticket — slowburn, chronicdegradationWithin budget — noaction

The exact multipliers and windows are yours to tune; 14x and 3x over 1-hour and 6-hour windows are a defensible starting point, not scripture. What matters is the shape: two severities, two regimes, both windows in agreement before anything fires.

Error budget burn-rate chart over a 28-day rolling window, with deploy, fast-burn alert, freeze, and thaw annotations

The hard part: wiring budget to a decision people honour

Everything up to this point is engineering, and engineering is the easy half. You can instrument honest SLIs, set a budget, and build burn-rate alerting that doesn't lie, and still have an SLO programme that changes nothing — because when the budget runs out, the release ships anyway. It ships because the feature was promised, because the quarter is ending, because a director wants it, and because "the error budget is exhausted" is not, on its own, a sentence that stops a release. It needs to have been made into one in advance.

The mechanism is a policy agreed before the budget is under pressure, in writing, by someone with the authority to bind both engineering and product. The canonical form is simple: when the error budget for a service is exhausted, all non-essential changes to that service freeze until the rolling window restores headroom. No new features. Reliability work and security fixes only. The release train stops.

The reason this has to be agreed in advance is the same reason you sign the contract before the dispute, not during it. In the calm of planning, "we'll freeze releases if we blow the budget" is an easy thing to nod along to. In the heat of a slipping deadline, with a budget at zero and a VP asking why the feature isn't out, it is an impossible thing to propose for the first time. The freeze only works if the argument was already won, on paper, when nobody had skin in the specific game. That document is not bureaucracy. It is political cover — the thing you point at so the decision is the policy's and not yours, which is the difference between enforcing a freeze and becoming the person who blocked the launch.

I have watched this fail in the predictable way more than once. The SLOs were defined, the budgets were sound, the alerting was textbook, and there was no signed policy — so the first budget exhaustion turned into a meeting, the meeting turned into an exception, the exception turned into precedent, and within two quarters the budget was a number on a dashboard nobody enforced. The technical work was immaculate. The programme died of a missing signature.

Make the consequence automatic where you can. A budget-exhaustion state that flips a flag the deployment pipeline actually reads — and refuses non-exempt changes against — removes the per-incident negotiation entirely. Nobody has to be the one who says no, because the system already said it, and overriding it requires a deliberate, logged, accountable act rather than a quiet word. The freeze you have to argue for each time is the freeze you will eventually stop arguing for.

When to break your own SLO

A policy you never override is a policy you haven't stress-tested, and a freeze that cannot bend will get torn out the first time it blocks something that genuinely had to ship. The competent move is not rigidity. It is a deliberate, recorded exception.

There are honest reasons to spend through a freeze. A regulatory deadline with legal consequences outranks an internal reliability target — the budget is a tool for serving the business, not a deity to be appeased while the business takes a fine. A security fix is reliability work by another name and should be exempt by default. And occasionally the budget itself is wrong: a flapping synthetic check or a miscalibrated SLI can exhaust a budget without any user feeling a thing, and freezing real work over a measurement artefact is its own kind of negligence.

The discipline is in how you break it. The override is explicit, it is logged, it names the person who authorised it, and — the part everyone skips — it triggers a review of why the budget was where it was. An exception that just waves the change through teaches the organisation that the freeze is theatre after all. An exception that costs a short, slightly uncomfortable retrospective keeps the policy credible, because everyone can see that breaking it is possible but not free. If breaking your SLO is costless, you do not have an SLO. You have a suggestion.

Reliability is a negotiation

Strip away the dashboards and the burn-rate maths and what remains is an agreement between the people who build a thing and the people who depend on it, about how often it is allowed to let them down. The number is the easy part. Picking an indicator that tracks real pain, setting a budget the bound parties actually agreed to, and writing down — in advance, with a signature that means something — what happens when the budget runs dry: that is the part that makes the difference between a programme that survives a quarter and a green tile nobody opens.

So write the terms down. Then the next time a release wants to ship into an exhausted budget, the answer is already on paper, signed by someone more senior than the argument, and your job is reduced to pointing at it. Which is a far better job than being the answer yourself.