Skip to main content
scaling.cloud
Back to the Guide

The quality signals that speed hides

A response can be fast and still be bad — silent for twenty minutes in the middle, acknowledged by someone who then vanished, technically resolved but never explained to the people who were waiting. The metrics that capture this live beyond the stopwatch. Learn to measure update discipline and genuine engagement, and why these signals catch failures that speed metrics are structurally blind to.

Imagine two incidents with identical timelines. Both were acknowledged in three minutes and resolved in fifty. By every speed metric they are twins. But in the first, the responder posted an update every ten minutes — what they were seeing, what they had ruled out, what they were trying next — so that everyone watching knew the situation was in competent hands. In the second, the responder acknowledged the page and then went silent for forty minutes, heads-down in the problem, while a support team fielded furious customers with nothing to tell them and a manager refreshed an empty channel wondering if anyone was even working on it. These are not the same incident. One was handled well and one was handled badly, and no measurement of speed can tell them apart. The difference lives entirely in the dimension that the stopwatch cannot see: the quality of the response as an act of communication and coordination, not just repair.

This is the case for measuring beyond speed. The fastest metrics are the easiest to collect and the easiest to understand, which is precisely why they crowd out the signals that matter just as much. Acknowledgment and resolution are events — a single timestamp each — and a response is not two events, it is everything that happens in between. A team that only measures the endpoints is flying blind through the middle of every incident, which is exactly where coordinated responses and chaotic ones diverge. The quality signals are harder to capture and less intuitive to report, but they are where the real maturity of a practice shows up, and a program that ignores them will keep producing responses that look good on a dashboard and feel terrible to live through.

Measure the longest silence, not the number of updates

The most important quality signal in any incident is whether the people depending on the response were kept informed, and the instinct is to measure this by counting updates: more posts, better communication. Resist that instinct, because update count is one of the most gameable numbers in the entire discipline. Tell a team to post more and you will get more posts — bursts of three trivial messages in a minute to satisfy a quota, padding that adds noise without adding information. Counting updates measures activity, not discipline, and the two come apart the moment the count becomes a target.

The signal that actually captures communication discipline is the opposite shape: not the total volume but the longest gap. For any incident, find the largest stretch of time between consecutive human touches on the timeline — the quietest moment, the longest the audience went without hearing anything. That single number is far harder to game, because the only way to shrink it is to genuinely communicate at a steady cadence; you cannot fix a long silence by burst-posting somewhere else. Bound the measurement by the incident's own start and end, so that an incident with no updates at all scores its entire duration as one unbroken silence rather than registering as a tidy zero. This matters enormously: the worst communication failures — total silence — must be the loudest in your data, never the most invisible. Aggregated across incidents as a median, the longest-silence signal tells you something speed never will: whether your responses keep people in the loop or leave them in the dark while the clock runs.

Watch for engagement, not just acknowledgment

The second quality signal addresses a failure that acknowledgment metrics actively conceal. Acknowledgment is supposed to mean "I have this," but in practice it often means only "I made the alert stop buzzing." A half-asleep responder who taps a notification to silence it, then rolls over, has acknowledged the incident by every measurement while doing nothing about it. The faster and more frictionless your acknowledgment mechanism, the more this gap can hide inside it: you optimize acknowledgment time beautifully and have no idea that some fraction of those acknowledgments were reflexes, not engagements.

The way to surface this is to measure the step after awareness — the gap between someone acknowledging an incident and the first real evidence that a human is actually working it: the first update they post, the first action they take, the first sign of genuine engagement rather than mere receipt. A healthy response shows that engagement following quickly behind acknowledgment. A troubling pattern shows acknowledgments that are not followed by anything for a long time, which tells you that your alerts are reaching people who cannot or will not act on them — perhaps because they are exhausted, perhaps because the page lacks the context to start, perhaps because it went to someone without the access to do anything. This measurement is especially revealing for incidents that your monitoring opened automatically, where there is no human at the moment of creation and the only question that matters is how long until a person genuinely showed up. The gap between the machine noticing and a human engaging is one of the purest measures of whether your on-call is staffed by people who are present, not just people who are paged.

Quality signals resist gaming because they measure outcomes, not actions

What these signals share is the property that makes them trustworthy: they are hard to fake without actually doing the thing they measure. You cannot shrink your longest silence without genuinely communicating throughout the incident. You cannot improve your engagement gap without responders who genuinely engage. Compare that to the easily-gamed alternatives — update counts, raw acknowledgment speed — and the difference is structural, not cosmetic. A good quality metric is one where the only path to a better number runs directly through better behavior, with no shortcut that satisfies the measurement while betraying its intent. Whenever you design a new signal, this is the test to apply: is there a cheap way to make this number look good that does not involve the team actually getting better? If there is, the metric will find it.

Speed metrics will always be the headline, because they are legible and because slowness genuinely matters. But a practice that measures only speed is measuring only the skeleton of a response. The flesh of it — whether people were kept informed, whether the person who acknowledged actually showed up, whether the long quiet stretches that terrify a waiting audience were avoided — lives in the quality signals. They are harder to collect and less satisfying to quote, and they are exactly the measurements that separate a response program that looks good from one that genuinely is.

Do this in scaling.cloudSet this up step by step in the docs.Open the how-to guide

Put this into practice.

scaling.cloud gives your team the on-call schedules, paging, and incident timelines this guide describes.