Reading the human cost of on-call
Every dashboard tells you how your systems are doing and almost none tell you how your people are doing — yet the sustainability of on-call is the constraint that decides whether your response practice survives. Learn the signals that reveal when a rotation is quietly burning people out, why out-of-hours load is the metric that matters most, and how to read the data before your best responders leave.
Most incident measurement is implicitly about the system: how fast it recovered, how often it broke, how well the response performed. This is natural, because the system is what the business is anxious about and the system is what generates the data. But it produces a strange blind spot. The single most important input to your incident response — the people who carry the pager — is the one thing the standard dashboards say almost nothing about. You can have flawless resolution times and a rotation that is six weeks from collapse, and nothing in your reliability reporting will warn you. The signals that matter for human sustainability are different from the ones that matter for system health, and a program that only watches the system will discover the human cost the way everyone eventually does: when a good engineer quietly hands back the pager and starts interviewing elsewhere.
The reason this matters beyond simple decency is that on-call health is not a soft concern competing with the hard ones; it is the constraint that determines whether all the other metrics are sustainable. A response practice runs on the willingness of experienced people to be woken up. Burn through that willingness and you do not get a slightly tired team — you get attrition of exactly the people whose judgment made your incidents go well, a rotation backfilled with responders who lack their context, and a slow, invisible decay in the quality of every number you do measure. The human cost is upstream of everything. Measuring it is not a wellness initiative; it is protecting the foundation the rest of your practice stands on.
Out-of-hours load is the signal that matters most
If you measure one thing about the human cost of on-call, measure how much of the disruption falls outside normal working hours. A page at two in the afternoon is an interruption; a page at two in the morning is a different category of harm — it shatters sleep, it spills into the next day's work, it accumulates as a debt the body actually keeps. The same number of incidents distributed across business hours versus scattered through the night represent wildly different costs to the humans absorbing them, and a raw incident count cannot tell the two apart. Out-of-hours load can. By classifying each page against the responder's actual working hours and tracking what fraction lands outside them, you convert a vague sense that "on-call has been rough lately" into a number you can watch, compare across rotations, and act on before it does lasting damage.
Defining "out of hours" demands a little care, because working hours are not universal. A responder following the sun from a different timezone has different normal hours than their colleague; a team with explicitly configured business hours has a precise answer while a team without one needs a sensible default. The honest approach classifies each page against the best available definition of the affected person's own hours — their configured schedule when one exists, a reasonable convention in their timezone when it does not — rather than against a single global office clock that misrepresents half the team. The metric will never be a perfect per-page audit; a multi-region responder may occasionally be misclassified, and that is acceptable, because the value of the signal is in the trend. A rotation whose out-of-hours fraction is climbing month over month is heading somewhere bad regardless of any individual edge case, and that trajectory is exactly what you need to see early.
Watch the distribution, not just the total
A second human-cost failure hides inside totals that look fine. A rotation can have a perfectly reasonable average load while quietly destroying one or two people, because the pages are not evenly distributed. The engineer who happens to own the flakiest service, or who is too conscientious to let a page go unanswered, or who keeps getting escalated to because everyone trusts them, can absorb a wildly disproportionate share of the pain while the team average stays comfortable. Measure only the total and you will congratulate yourself on a healthy rotation right up until your most reliable person breaks. The distribution is the metric; the total is a comforting average that hides the people the rotation is actually hurting.
This means looking at per-person load and especially at the worst cases — who is taking the most pages, who is taking the most night pages, who is being escalated to most often outside their own primary shifts. Fairness in on-call is not a nicety; it is a sustainability requirement, because an unfair distribution concentrates burnout in your most valuable people. When you find the load skewed, the response is rarely to ask that person to be more resilient. It is to ask why the work is landing there: a service that needs fixing, an escalation policy that overuses one person as a safety net, a knowledge silo that makes one engineer irreplaceable in a way that is dangerous rather than flattering. The metric points at a person, but the fix is almost always in the system around them.
The data tells you to fix the cause, not push the people
The trap in human-cost metrics is identical to the trap in every other incident metric, just with higher stakes: it is tempting to treat the number as a target to be driven down by pressuring the people, when its only legitimate use is to point at causes in the system. A high out-of-hours load does not mean responders should toughen up; it means something is generating too many nighttime pages and that something — a noisy alert, a fragile dependency, a service that fails on a schedule — is the actual problem. An uneven distribution does not mean the overloaded person should complain less; it means the work needs rebalancing or the underlying fragility needs fixing. Read the human-cost data this way and it becomes one of the most powerful prioritization tools you have: it tells you which reliability work to do first by showing you which fragility is costing your people the most sleep.
The deepest reason to measure the human cost is that it is the one part of your practice that will not show up in a postmortem or a system dashboard until it is too late to fix gently. A failing service announces itself. An exhausted rotation does not — it stays quiet and functional and uncomplaining right until the people in it leave, and by then the cost is irreversible. Watching out-of-hours load and the fairness of the distribution is how you hear the warning while you can still respond to it. It is the measurement that keeps your response practice from succeeding its way into collapse.