Insights
You cannot improve what you refuse to look at, and you cannot look honestly at incident response without numbers — but the wrong numbers are worse than none, because they create the comfortable illusion of progress while the real work quietly decays. Insights is the discipline of measuring incident response in a way that makes teams better rather than busier or more frightened: knowing what the headline speed metrics actually tell you and where they lie, watching the quality signals that speed hides, reading the human cost a dashboard rarely shows, and turning all of it into changes to the system instead of pressure on people. These sections cover the metrics that matter and their limits, the response-quality signals beyond speed, the data that tells you when on-call is hurting your team, and how to run a measurement practice that improves the process without gaming it.
In this chapter
- The metrics that matter, and where they lieMTTA and MTTR are the two numbers every incident program eventually reports, and for good reason — they are the closest thing the discipline has to a universal vocabulary for speed. But a number you do not understand is a number that will mislead you. Learn what these metrics genuinely measure, the specific ways they distort, and how to read them so they sharpen your judgment instead of replacing it.
- The quality signals that speed hidesA response can be fast and still be bad — silent for twenty minutes in the middle, acknowledged by someone who then vanished, technically resolved but never explained to the people who were waiting. The metrics that capture this live beyond the stopwatch. Learn to measure update discipline and genuine engagement, and why these signals catch failures that speed metrics are structurally blind to.
- Reading the human cost of on-callEvery dashboard tells you how your systems are doing and almost none tell you how your people are doing — yet the sustainability of on-call is the constraint that decides whether your response practice survives. Learn the signals that reveal when a rotation is quietly burning people out, why out-of-hours load is the metric that matters most, and how to read the data before your best responders leave.
- Using data to improve, without gaming itThe moment a metric becomes a target, people start optimizing the number instead of the thing it was supposed to represent — and incident response is unusually easy to game. Learn how to run a measurement practice that actually drives improvement: which metrics to expose to whom, why trends beat thresholds, and how to keep the numbers honest by never turning them into a scoreboard.