Skip to main content
scaling.cloud
Back to the Guide

Writing incident updates people trust

An incident update is a promise about cadence, a statement of impact, and a small act of honesty all at once. Learn to write updates that tell customers what they actually need — what's broken, who it affects, and when they'll hear from you next — without overclaiming, hedging, or going silent.

The instinct during an incident is to say nothing until you have something definitive to say. It feels responsible — why broadcast confusion, why commit to a cause you haven't confirmed, why promise a fix time you might miss? But silence is not neutral. To the customer staring at a failed checkout or a frozen dashboard, silence reads as one of two things: either you don't know yet, or you don't care. Both are worse than an honest "we're aware of this and actively working on it." The first rule of incident communication is that the absence of an update is itself a message, and almost always the wrong one.

A good incident update does a surprisingly small number of jobs. It tells people what is affected, in terms they can map to their own experience. It tells them who is affected, so the unaffected majority don't panic and the affected minority feel seen. It tells them what you're doing about it, at a level of detail that conveys competence without leaking noise. And — the part most teams forget — it tells them when they'll hear from you next. Get those four things right and you have an update that does its job. Everything else is decoration.

Cadence is a promise, and silence breaks it

The single most reassuring sentence in any incident update is the one that says when the next one is coming. "We'll post another update within 30 minutes" turns an anxious refresh-the-page loop into a calm wait. It tells the reader that someone is actively shepherding this, that they won't be left guessing, and that the channel they're watching will keep producing signal. It also, crucially, binds you: having promised an update in 30 minutes, you now have to produce one, even if the only honest content is "still investigating, no change yet, next update in 30 minutes."

That "no change" update feels pointless to write and is anything but. Customers do not lose faith because an incident is taking a while; they lose faith because they were promised a follow-up that never came. A heartbeat update that says nothing new still says the most important thing — we are still here, still on it, and still keeping you informed. Pick a cadence that matches severity (tighter for a full outage, looser for a degraded-but-usable state), state it explicitly, and then honour it religiously. Missing a self-imposed update deadline does more reputational damage than the outage itself.

Match the altitude to the audience

The people reading a public incident update are not your engineers, and writing for the wrong audience is how good information becomes useless. A customer does not need to know that a connection pool exhausted or that a canary deploy triggered a cache stampede. They need to know that logins are failing for some users in Europe and that you're working on it. The technical detail isn't wrong, it's just aimed at the wrong reader — it raises questions it can't answer and makes the unaffected wonder if they should be worried.

The discipline here is to describe impact, not cause. "Some users may be unable to upload files" is a sentence anyone can act on: an affected user knows to wait, an unaffected one knows to carry on. "We are rolling back a recent deployment to the storage service" is engineering chatter that belongs in your internal channel. Lead every external update with the symptom as the customer experiences it, scope it as precisely as you honestly can, and only then — if at all — gesture at what you're doing. When you genuinely don't know the scope yet, say that, plainly: "We're still determining how many users are affected." Naming your uncertainty is far more credible than a confident guess that turns out wrong.

Honesty compounds; spin decays

There is a strong temptation, under pressure, to soften. To call a total outage "degraded performance," to say "a small number of users" when it's most of them, to imply you're closer to a fix than you are. Every one of these is a loan against your future credibility at a punishing interest rate. Customers experience the real impact directly; if your words and their reality diverge, they don't conclude they misread — they conclude you're managing them. And once a customer decides your status updates are PR rather than information, every future update is discounted to zero, including the honest ones.

The opposite habit — calibrated, slightly conservative honesty — pays out the other way. Describe impact at least as severe as it actually is. If you're not sure whether something is fixed, say "we believe this is resolved and are monitoring" rather than declaring victory and risking a humiliating reopen. Acknowledge the disruption plainly, without grovelling and without minimising: "We know this is disruptive and we're sorry" lands; three paragraphs of corporate contrition do not. The teams that come out of a bad incident with their reputation intact or improved are almost always the ones whose communication was so straight that customers trusted them more afterward than before.

Write the resolution like it matters, because it does

The final update is the one people screenshot, forward, and remember, and it deserves more care than the heat-of-the-moment ones. A good resolution message confirms unambiguously that things are working again, states what the impact was and over what window (so customers can reconcile it with what they saw), and — without turning into a full postmortem — gives a one-line sense of what happened and what you're doing to prevent a repeat. "Between 14:05 and 14:50 UTC, roughly a third of users were unable to complete checkout due to a database failover that didn't promote cleanly; we've corrected the failover configuration and are reviewing our promotion automation." That single sentence converts a frustrating experience into evidence that you understand your own system and take it seriously.

What the resolution should not do is overpromise on the future. Don't commit to a fix that doesn't exist yet or a guarantee you can't keep; "we're investigating how to prevent this class of failure" is honest, "this will never happen again" is a hostage to fortune. Close the loop, own the impact, point at the next step, and stop. The whole arc — first acknowledgment, steady heartbeats, calibrated honesty, a clean resolution — is what turns the worst hour of your week into the thing a customer cites when they tell someone else they trust you. Communication during an incident is not damage control bolted onto the real work; for everyone outside the war room, it is the work.

Do this in scaling.cloudSet this up step by step in the docs.Open the how-to guide

Put this into practice.

scaling.cloud gives your team the on-call schedules, paging, and incident timelines this guide describes.