Skip to main content
scaling.cloud
Back to the Guide

Learning

The work of an incident is not finished when the service recovers — that is when the most valuable work begins. Learning is the discipline of turning an outage into something the whole organization is better for having survived: a blameless post-mortem that explains what really happened, follow-ups that actually get done instead of decaying in a backlog, and a review culture that rehearses for failure before it arrives. These sections cover how to run a post-mortem people read, how to write follow-ups that change the system rather than blame a person, and how to build the muscle of learning into a team so every incident leaves it stronger.

In this chapter