Back to the Guide
Incident response
Incident response is what happens between "something is wrong" and "we are back to normal" — and whether that interval is ten calm minutes or three chaotic hours is mostly a matter of process, not heroics. These sections cover what counts as an incident and why declaring early wins, severity that ranks impact honestly, the response roles that keep a busy room coordinated, and the lifecycle that carries an incident from first signal to resolution.
In this chapter
- What counts as an incident, and why to declare earlyThe cheapest mistake in incident response is declaring something that turns out to be nothing. The expensive one is the silent hour where everyone assumed someone else had it. Learn what actually warrants an incident and why the bias should always be toward declaring sooner.
- Severity levels that actually rank impactA severity scale exists to answer one question under pressure — how hard do we push, and who do we wake? Learn to build levels defined by impact rather than guesswork, anchored to observable signals, and free of the urgency-versus-importance confusion that breaks most scales.
- Response roles that keep a busy room coordinatedWhen a serious incident pulls in half the team, the failure mode is no longer too few hands — it's too many, all uncoordinated. Learn the small set of roles that turns a crowd into a response: a single accountable lead, a dedicated comms voice, and clean handoffs.
- The incident lifecycle, from triage to resolutionEvery incident moves through the same shape — get acknowledged, get understood, get mitigated, get resolved — even though the technical work is different every time. Learn the lifecycle that gives a chaotic response a backbone, and why mitigation and resolution are not the same milestone.