The Pager Is a Smoke Detector, Not a Doorbell: Designing On-Call That Doesn't Burn People Out

Your pager went off at 02:14 this morning. By the time you had found your laptop, joined the bridge, and read the alert, it had resolved itself. The cause was a node that briefly went unhealthy and then came back, as nodes do. Nobody did anything, because there was nothing to do, and there was nothing to do because the alert was never about a thing anyone could act on. It was about a thing that happened. You were woken to be informed.
That is the resting state of most on-call rotations, and it is worth being precise about why it is intolerable. The problem is not that the pager fires. The pager is supposed to fire. The problem is the doorbell — the alert that goes off because something occurred, regardless of whether anyone needs to come to the door. A smoke detector earns its place on the wall by staying silent for years and then screaming exactly once, at the moment it matters, in a way you cannot ignore. A doorbell rings whenever a thing arrives, useful or not, and after enough false arrivals you stop hearing it. Most monitoring stacks are walls of doorbells, and the rotation has long since learned not to hear them.
The doorbell problem
Cause-based alerting is the doorbell, and it is the default because causes are easy to measure. CPU is at 90%. The queue has 10,000 messages in it. A pod restarted. Disk is 85% full. These are real facts about the system, they are sitting right there in your metrics already labelled, and turning each one into an alert feels like diligence. It is the opposite. Each of those facts may or may not correspond to a user feeling anything, and the alert has no way of knowing which. CPU at 90% on a service that is comfortably serving every request inside its latency target is not an incident. It is a Tuesday. Paging someone for it teaches them that the pager lies.
The cost of that lesson compounds. Every alert that fires without a corresponding action erodes the signal value of the channel it arrives on, and the erosion is not linear — it is a phase transition. Below some noise threshold the rotation triages every page seriously; above it, the rotation starts pattern-matching alerts to "probably nothing" before reading them, and the moment they do that, the real page is just as likely to be dismissed as the noise. Alert fatigue is not tiredness. It is a learned, rational response to a channel that has cried wolf, and the rational response gets people hurt — because the one night the wolf is real, it arrives on the same channel, at the same hour, in the same tone as the four hundred nights it wasn't.
And the people leave. Pager-driven attrition is the quiet line item nobody costs properly. The engineer who gets woken twice a week for nothing does not file a complaint; they update their CV. You lose your most operationally experienced people first, because they have been on the rotation longest and remember most clearly what it has cost them. The replacement cost of a senior SRE dwarfs the engineering time it would have taken to delete the alerts that drove them out — but we rarely put those two numbers on the same page, which is precisely why the doorbells stay on the wall.
Symptom over cause, page on burn
The fix is to stop alerting on what the system did and start alerting on what the user felt. Alert on symptoms, not causes. The user does not care that CPU is at 90%; they care that checkout is slow. If CPU is at 90% and checkout is fast, there is no symptom and there should be no page. If checkout is slow, that is a symptom, and it warrants attention whether the cause is CPU, a slow dependency, a bad deploy, or something you have never seen before. The symptom-based alert catches the failure modes you did not anticipate, which is most of them. The cause-based alert only catches the ones you did, and pages you for a thousand non-events in between.
Symptom-based alerting connects directly to the error budgets I wrote about a few days ago. If you have an SLO, you already have the only symptom that matters expressed as a number: the rate at which you are burning the budget. Page on burn rate, not on raw metrics. A burn rate of 1 is sustainable by definition — you will exhaust the budget exactly at the end of the window, which is what the budget is for. A burn rate of 14x over an hour means you will blow a month's budget by breakfast; that is a smoke detector going off, and it warrants waking someone. A burn rate of 3x sustained over six hours is a slow leak — real, worth fixing, and emphatically not worth anyone's sleep. It gets a ticket and ruins someone's Tuesday in daylight, which is the humane way to lose an argument with entropy.
Everything else — the node that flapped, the queue that drained on its own, the pod that restarted and recovered — was never a page. It is information, and information that requires no immediate human action belongs in a ticket or on a dashboard, not on a pager at 02:14. The discipline is in routing each signal to the channel that matches its urgency, and the test for "page" is brutally simple: would you wake a competent, tired colleague for this, knowing what it is? If the honest answer is no, it is not a page. The smoke-detector test does most of the triage on its own.
| Signal | Page now | Ticket | Dashboard |
|---|---|---|---|
| Checkout journey error rate burning budget at 14x over 1h | ✓ | ||
| Latency creeping, burning budget at 3x over 6h | ✓ | ||
| Single node unhealthy, service unaffected | ✓ | ||
| Disk 85% full, growing slowly, days of headroom | ✓ | ||
| Disk 98% full, hours to exhaustion, no auto-remediation | ✓ | ||
| Payment provider returning 5xx, users cannot pay | ✓ | ||
| Background job 2% slower after a deploy, within freshness SLO | ✓ | ||
| Certificate expires in 21 days | ✓ | ||
| Certificate expires in 6 hours | ✓ |
The pattern in that table is not the specific thresholds; those are yours to tune. It is that the right-hand columns are where most of your current pages belong, and the act of moving them there is the entire job.
Every page needs an owner, an action, and a runbook
A page that survives the triage above still has to earn its place once it fires. The bar is three things: it has a clear owner, it has an action that owner can take now, and it has a runbook link in the alert payload that tells them what that action is. An alert that pages a generic rotation with no runbook and no obvious next step is just a more stressful way of telling someone something is wrong. The point of waking a human is that a human can do something; if the only available action is to acknowledge the page and watch a graph, you have built an expensive, sleep-depriving dashboard.
The runbook link is not optional and it is not documentation for later. It goes in the alert itself, so the half-awake responder is one click from "here is what this means and here is what to do", not one frantic Confluence search from it. If you cannot write that runbook — if you genuinely do not know what someone should do when this fires — that is not a gap in your documentation. It is proof the alert should not page, because you have just admitted there is no action. Delete it, or demote it to a ticket where someone can work out the action in daylight.
Rotations that respect sleep
The structure of the rotation matters as much as the alerts that feed it. A one-person rotation with no secondary is a single point of failure made of a human being, and humans fail in the specific way of being asleep, in a tunnel, or simply broken by the third 03:00 wake-up that week. You need a primary, a secondary who is paged when the primary does not acknowledge inside a few minutes, and an escalation path that reaches a human with authority before it reaches nobody at all. Tiers and time windows are not bureaucracy; they are the difference between an incident that gets handled and one that pages into a void until morning.
Where you have the geography for it, follow-the-sun is the most humane structure there is, because it does the one thing no amount of alert-tuning can: it stops paging people in the middle of their night. A rotation split across three regions means every page lands on someone awake, at their desk, and paid to be there. Most teams do not have three regions, and for them the honest move is to cap the rotation length, guarantee recovery time after a bad night, and treat a week of broken sleep as the operational cost it is rather than a badge of seniority.
Measure it. Track page rate per person per week and interrupt rate — pages outside working hours — as first-class operational metrics, reviewed as seriously as latency. A rotation trending upward on either is a fire you can see from across the room, and the budget you are spending is human. Then, having measured it, spend a month deleting alerts: take every page from the last quarter, and for each one ask whether anyone took an action, and whether that action could not have waited until morning. The alerts that fail both tests go. Deleting an alert feels like negligence and is the opposite — it is the single highest-leverage reliability work most teams never schedule, because a deleted alert protects every future night the rotation would otherwise have lost to it.
A quiet pager is not a sign that nobody is watching. It is the sign that the watching is working — that the wall holds one smoke detector that means it, and no doorbells at all.