Using data to improve, without gaming it
The moment a metric becomes a target, people start optimizing the number instead of the thing it was supposed to represent — and incident response is unusually easy to game. Learn how to run a measurement practice that actually drives improvement: which metrics to expose to whom, why trends beat thresholds, and how to keep the numbers honest by never turning them into a scoreboard.
There is a law that everyone who measures anything eventually learns the hard way: when a measure becomes a target, it stops being a good measure. The instant a number is used to judge people, they begin optimizing the number rather than the reality it was meant to reflect, and they are usually clever enough to do it without technically lying. Incident response is unusually vulnerable to this because almost every one of its metrics has a cheap shortcut. Want a better resolution time? Stop declaring small incidents, or close them early and reopen quietly. Better acknowledgment time? Acknowledge reflexively before you have even read the page. Better communication score? Post three empty updates to satisfy the cadence. None of these improves a single real outcome, and all of them will appear on your dashboard as progress. A measurement practice that does not actively guard against this will, over time, produce numbers that look better every quarter while the actual practice they describe slowly gets worse.
The goal of insights is improvement, and improvement is a genuinely good thing to want. But the path from measurement to improvement runs straight through the gaming problem, and a program that does not think carefully about incentives will find that its data has quietly detached from reality. Getting this right is less about which metrics you choose than about how you use them — who sees them, what conversations they feed, and whether they are ever allowed to become a scoreboard.
Trends beat thresholds, and direction beats absolutes
The first discipline is to care about movement, not magnitudes. An absolute target — "all incidents resolved within thirty minutes" — invites gaming almost automatically, because the target is a line and there is always a way to get on the right side of a line without deserving to. A trend asks a different and far more useful question: are we getting better or worse, and at what rate? Direction is much harder to fake than position, because faking a sustained improvement trend means continuously gaming the metric in the same way, which is more effort than actually improving and tends to get noticed. Watching trends also respects the reality that the right number is contextual: a thirty-minute median might be excellent for one class of system and alarming for another, and a single threshold imposed across both will be wrong for at least one of them. The question is rarely "are we below the line" and almost always "are we moving in the right direction, and is the rate of change something we are happy with."
Trends also protect you from the most common misreading of incident data, which is mistaking noise for signal. Incident metrics are volatile by nature — one bad month with a single nasty outage can move any number — and a practice that reacts to every monthly wobble will exhaust itself chasing ghosts and will train its people to fear the metric. Looking at the trend over a meaningful window filters that noise and keeps attention on the changes that are real. A metric that jumped this month and a metric that has been deteriorating for two quarters demand completely different responses, and only the trend view tells you which you are looking at.
Expose metrics to learn, never to rank
The single most consequential decision in a measurement practice is what the numbers are allowed to be used for. There is a bright line between using metrics to understand the system and using them to rank people, and crossing it is the fastest way to poison your data. The moment an engineer believes that their personal acknowledgment time or their incident count will be compared against a colleague's in a performance conversation, the metric stops measuring response quality and starts measuring their willingness to be seen taking incidents. They will decline to acknowledge things, avoid declaring incidents, and route pain away from their own name — and your data will degrade in exactly the dimensions you most wanted it to be honest about. Aggregate metrics that illuminate the system are a gift; individual metrics that rank people are a trap, and the same number can be either depending entirely on how you frame and deploy it.
This does not mean individual data is useless — the human-cost signals depend on it, and spotting that one person carries too much of the load requires looking at people specifically. The distinction is in purpose, not in granularity. Looking at an individual's load to find an unfair distribution you should fix is using the data to improve the system. Looking at an individual's load to decide who is pulling their weight is using it to rank, and the people being measured will instantly tell the difference, because their careers depend on telling the difference. Keep the metrics in the retrospective and the planning conversation, where the subject is "what should we change about the system," and keep them out of the performance review, where the subject is "how is this person doing." The same dashboard is a tool of improvement in one room and a weapon in the other.
Let the numbers ask questions, and let humans answer them
The healthiest way to think about an insights practice is that its metrics are not answers but questions. A rising longest-silence median does not tell you that your team communicates badly; it tells you to go look at why the silences are growing, and the answer might be a tooling gap, a staffing problem, or a single recurring incident type that is hard to narrate. A worsening tail in resolution time does not tell you responders got slower; it tells you to investigate what is in the tail, and you might find one gnarly class of failure that needs an architectural fix rather than any change to how people respond. Every metric movement is the start of an inquiry, not the end of one, and the inquiry is where the actual improvement comes from. A team that treats numbers as prompts for human investigation gets better; a team that treats them as verdicts to be optimized gets better numbers and a worse practice.
This is why the most valuable place for incident data is the retrospective, fed back to the team that lived the incidents, framed as "here is what the data is showing us — what do we think is going on?" Used that way, measurement closes the loop that makes a response practice mature: incidents generate data, data surfaces patterns, patterns drive investigation, investigation drives changes to the system, and the changed system produces better incidents whose data you measure again. The numbers are the instrument that makes the loop visible, but the improvement itself is always human — a person noticing a pattern, asking why, and fixing a cause. Keep the metrics in service of that human judgment, never above it, and they will make your practice honestly better. Turn them into targets and scoreboards, and they will make your dashboards better while the practice they were supposed to improve quietly hollows out.