Back to the Guide
Learning
The work of an incident is not finished when the service recovers — that is when the most valuable work begins. Learning is the discipline of turning an outage into something the whole organization is better for having survived: a blameless post-mortem that explains what really happened, follow-ups that actually get done instead of decaying in a backlog, and a review culture that rehearses for failure before it arrives. These sections cover how to run a post-mortem people read, how to write follow-ups that change the system rather than blame a person, and how to build the muscle of learning into a team so every incident leaves it stronger.
In this chapter
- Running a blameless post-mortemA post-mortem is not a trial, and it is not a formality. It is the one moment when a team can look honestly at how its system actually behaved under stress. Learn what makes a post-mortem blameless in practice, not just in name, and why that is the only kind worth running.
- Writing post-mortems people actually readA perfect investigation buried in a document nobody opens has changed nothing. The hard part of a post-mortem is not finding the truth — it is writing the truth so that a busy stranger reads it, believes it, and acts on it. Learn to write for the reader you will never meet.
- Turning learnings into follow-ups that get doneThe most common failure in post-incident learning is not a bad investigation — it is a good one whose follow-ups quietly die in a backlog. Learn how to write follow-ups that change the system, give them an owner and a deadline, and make sure they actually ship before the next incident proves they should have.
- Building a culture that learns from failurePost-mortems and follow-ups are practices; a learning culture is the thing that makes them stick when no incident is forcing the issue. Learn how to make learning a habit rather than a reaction — through blameless norms, rehearsed failure, and shared knowledge that outlives any individual.