Skip to main content
scaling.cloud
Back to the Guide

Severity levels that actually rank impact

A severity scale exists to answer one question under pressure — how hard do we push, and who do we wake? Learn to build levels defined by impact rather than guesswork, anchored to observable signals, and free of the urgency-versus-importance confusion that breaks most scales.

The first real decision in any incident is also one of the hardest: how bad is this? You make it with incomplete information, often half-awake, under the distorting pressure of the moment — and yet everything downstream hangs on it. Severity is the answer to that question, compressed into a single label that the rest of the response can act on without re-litigating. Done well, it lets a tired responder make a fast, defensible call about how hard to push and who to pull in. Done badly, it becomes a source of argument exactly when you have no minutes to spare for arguing.

A severity scale is, at bottom, a shared agreement about what different levels of "bad" mean and what each one obliges you to do. Its job is not to be precise — the real world does not sort itself into tidy tiers — but to be fast and consistent: to get two different people, on two different nights, to classify the same situation roughly the same way, and to know roughly the same things follow from that classification.

Rank by impact, not by cause or guesswork

The cardinal rule is that severity describes impact on the people who depend on your service, not the technical nature of the problem and not how stressed you feel. "A database is down" is not a severity — it is a cause. The severity is whatever that database being down does to your users: if it takes checkout offline for everyone, that is your top level; if it only slows an internal admin report nobody reads after hours, it is near the bottom, the very same root cause notwithstanding. Anchoring on impact keeps the scale honest, because impact is the thing your users and your business actually feel, and it is assessable even while the cause is still a mystery.

This matters most at the extremes, where instinct misleads. A dramatic-looking failure — alarms everywhere, a scary stack trace — can have trivial user impact, and a boring-looking one — a slow creep of elevated error rates — can be quietly catastrophic. If you rank by how alarming the symptom looks rather than how much the user hurts, you will routinely over-respond to the spectacular and under-respond to the silent, which is exactly backward. Force every severity judgment through the same question: who is affected, how badly, and how widely?

Keep the levels few and the meanings sharp

A scale with too many levels is a scale nobody can apply consistently, because the boundaries between adjacent tiers blur into a matter of taste. Four levels is a well-worn sweet spot — enough resolution to distinguish "drop everything" from "handle it this week," few enough that the line between any two is crisp. A discipline-neutral reading of four levels runs roughly:

  • Critical — severe, broad impact on core functionality; the business is visibly hurting and the response should pull in whoever it takes, immediately, including out of hours.
  • High — significant impact on an important capability, but bounded — a major feature degraded, or a smaller slice fully down. Worth an urgent, focused response, though perhaps not the whole company.
  • Medium — real but contained: a non-core feature impaired, a workaround exists, or only a small population is affected. Handle promptly, in working hours, without heroics.
  • Low — minor or cosmetic; noticeable to someone, but no meaningful disruption. Tracked and fixed on a normal cadence.

The exact wording matters less than two properties: the levels must be few enough that the distinctions stay sharp, and each must come with a clear sense of what it obliges — how urgently you respond and how far you escalate. A label that implies no action is decoration.

Anchor each level to observable signals

Abstract definitions drift. "Significant impact" means one thing to the engineer who built the feature and another to the one who's never touched it, and at 3am neither is in a mood to debate semantics. The fix is to anchor each level to concrete, observable anchors — the kind of thing you can check on a dashboard or state plainly, not infer from feeling. Tie your levels to questions with checkable answers: Is core functionality affected, or a peripheral feature? Are all users hit, a region, or a handful? Is there a workaround? Is revenue or safety or data integrity on the line?

The point of the anchors is to make the scale self-service. A responder should be able to look at a small set of concrete questions, answer them honestly from what they can observe, and arrive at a level without needing to summon a committee. That is what turns severity from a topic of debate into a quick, repeatable classification — and a quick classification is the whole reason the scale exists.

Severity ranks impact, and stays re-assessable

Two clarifications save more grief than any other. First, resist collapsing severity into a generic "urgency" or ordering knob. Severity answers how much impact; that is a different axis from what we choose to work on first, which also weighs effort, blast radius, and what else is on fire. Keeping severity purely about impact stops it from quietly becoming a catch-all ordering field that means everything and therefore nothing — let it rank impact, and let your response process decide what that impact warrants.

Second, severity is a live assessment, not a verdict carved at declaration time. You set it early on partial information, precisely so you can act; as the picture sharpens, you move it. Something declared at a high level that turns out to be one flaky node gets downgraded without ceremony or embarrassment; something that started as a contained blip and is now eating every region gets upgraded the instant you see it spreading. A scale people feel free to revise stays accurate, and an accurate severity is what lets everything downstream — who gets paged, what the public sees, how hard the room pushes — stay calibrated to the truth of the moment rather than to a guess made in its first, most ignorant minute.

Put this into practice.

scaling.cloud gives your team the on-call schedules, paging, and incident timelines this guide describes.