Most on-call rotations are not failing because the systems are unreliable. They are failing because the rotation pages too often, for things nobody can act on, and the people in it stop trusting the channel.
Google’s SRE practice is the most cited work on this, and the two rules that matter most in it are simple enough to adopt without a reliability department.
Rule one: page on symptoms
Alert on what a user experiences. Latency at a percentile, error rate as a proportion of traffic, saturation of whatever resource is closest to exhaustion, and traffic so the other three can be interpreted.
Do not page on causes. A full disk is a cause. A restarted pod is a cause. A failed node is a cause. Those belong on a dashboard, and they belong in the investigation that follows a symptom page.
The reason is not purity. It is that cause alerts multiply with the estate while symptom alerts stay roughly constant, and they fire during events where nothing is actually wrong for anybody.
If a user cannot tell, it is not a page. It is a graph.
Rule two: every page needs a human decision
The second test is harder and removes more alerts than the first.
If the response to an alert is always the same action, that action should be automated and the alert deleted. If the response is always “acknowledge and wait”, the alert is telling you the system recovered on its own, which is a dashboard entry. If the response is “look at it in the morning”, it is a ticket.
What is left is the small set where a person has to weigh something: declare an incident, fail over, roll back, call a vendor, wake someone else. That set is what a rotation should exist for.
- Page if
- a user notices and a human must decide
- Target
- a quiet week should be genuinely quiet
- Reviewed
- weekly, with deletions
This diagram as text
- What kind of signal is it?
- Symptom, user-visible — and a human must decide
- Cause — useful in the investigation
- Recurring — same action every time
- Informational
- Only the first wakes a person
- Page — someone is woken
- Dashboard
- Automate the action, then delete the alert
- Log, with retention
Relationships
- Symptom, user-visible → Page
- Cause → Dashboard
- Recurring → Automate the action, then delete the alert
- Informational → Log, with retention
The weekly review is the whole practice
One hour, once a week, with the person coming off call. Go through every page. For each one:
- Did a user notice?
- Did it need a human?
- What was done?
- What would stop it happening again, or stop it paging?
Then act on the answer. Delete the alerts that led nowhere. Automate the ones with a constant response. Fix the one underlying cause that produced the most pages.
An organisation that does this for three months finds its page volume falls by most of itself, and almost none of that comes from the systems getting better. It comes from the alerts getting honest.
The things that burn people out
Ranked roughly by how often we find them:
Alerts with no runbook. Being woken to a message you do not understand, for a system you did not build, with no written procedure, is the fastest route to someone leaving a rotation.
No handover. A week ends and the next person starts without knowing what is still open, what was worked around rather than fixed, or what is likely to fire again.
Pages for things the responder cannot fix. If the answer is always to call a vendor, or to wake the one person who understands that subsystem, then the rotation is a relay and the real on-call is whoever is at the end of it.
No compensation or no time back. Whatever the arrangement is, it should be explicit. Unstated expectations are where resentment accumulates.
A rotation too short to recover from. One week in three, in a noisy estate, with nothing done about the noise, consumes people predictably.
What to put in place first
If you are introducing on-call, or repairing one, do these in order:
- Name the person per period, in writing, with an escalation path and the hours it covers.
- Cut the alert set to symptoms that pass both tests. This usually removes most of them.
- Write a runbook for every remaining alert. If an alert has no runbook, either write one or accept that it is not a page.
- Add a handover, five minutes, written, at the end of each period.
- Start the weekly review and actually delete things in it.
None of this needs a product. It needs somebody with the authority to delete an alert.
It is also the evidence
ISO/IEC 27001 Annex A 5.24 asks for planned incident management: defined responsibilities, defined escalation, and evidence that it operated. A named rotation, a handover record and a weekly review with dated outcomes produce that as a by-product, which means the practice that makes on-call survivable is also the one that satisfies the auditor.
How STP approaches this
Where we run a rotation, the alert set passes both tests before the rotation starts, every alert has a runbook written to be executed by someone who did not build the system, and the weekly review happens with the authority to delete. Where we are reviewing an existing rotation, we start by counting the pages that produced no action, which is usually most of them and is usually the fastest thing to fix.
More on managed IT and monitoring and help desk and support, or start a conversation.

