Incident Communication Templates You Can Copy-Paste at 3 A.M.
Good incident communication is written before the incident. At 3 a.m., with the database down and three people in a call, nobody produces calm, precise customer-facing prose from a blank page — they either go silent (the worst option) or write something panicked they'll regret. The fix is boring and effective: templates for every stage, agreed on in daylight, filled in under pressure.
Here's the full set — first update, progress, resolution, follow-up — plus the tone rules that make them work and the internal/public split that keeps your war room honest.
The first update: acknowledge within minutes
The first update has one job: kill the silence. Customers hitting errors are asking is it just me? and do they know? — and every minute without an answer converts into support tickets and churn-flavored doubt. Aim to post within 5–10 minutes of confirming impact, on your status page first, everywhere else second.
Include:
- That you know, and you're on it
- What's affected, in customer language ("API requests," "dashboard logins" — not service names)
- Rough scope if you know it ("some users," "all requests to…")
- When the next update is coming — and keep that promise
Omit:
- Root-cause speculation ("we believe a database failover…") — early guesses are usually wrong, and wrong guesses erode trust when corrected
- ETAs you can't stand behind
- Blame — of a vendor, a deploy, or a person
The acknowledged-but-vague update beats the precise-but-late one, every time.
Progress updates: cadence is the message
Once the first post is up, the cadence is the communication. State it explicitly ("next update by 14:30 UTC") and hit it even when nothing has changed — "still investigating, no change" is information; silence is abandonment. As a default: every 30 minutes for a full outage, every 60 for partial degradation. Move through explicit stages so progress is scannable: Investigating → Identified → Monitoring → Resolved.
Resolution and follow-up
Resolve only when customers are recovered, not when the fix is deployed — that's what a Monitoring stage is for. The resolution note should say what happened at one paragraph of depth, the impact window in UTC, and whether any customer action is needed.
For significant incidents, follow up within a few days with a short public post-mortem: timeline, root cause in plain language, and what you're changing. Teams consistently find these build trust rather than damage it — the companies famous for great incident handling are famous because of their write-ups, not despite them.
Tone rules
- No blame. Not "our hosting provider failed" — customers bought from you. "A component of our infrastructure" is honest without finger-pointing. Internally, blamelessness is what keeps engineers reporting mistakes early.
- No over-promising. Never "resolved shortly," no ETA you're less than 90% sure of. The only safe promise is the next update time.
- Plain language. "Logins are failing," not "elevated error rates on the authentication layer."
- Calm, not casual. No exclamation marks, no "oops!", no memes. Warmth belongs in the apology, not the diagnosis.
- Own it. "We're sorry — this is on us" reads better in every follow-up than a passive "service interruptions were experienced."
Copy-paste templates
Swap the {placeholders}, keep the structure.
Investigating — the first post
Investigating — {Short, plain title, e.g. "API requests failing"} We're aware that {affected surface} is {failing / degraded} for {scope, e.g. "some users" / "all requests"} and are actively investigating. {If true: "Your data is not affected."} Next update by {time UTC}.
Identified
Identified We've identified the cause: {one plain-language sentence, no blame, e.g. "a configuration change affecting our database connections"}. A fix is {being prepared / rolling out now}. {Affected surface} remains {status}. Next update by {time UTC}.
Monitoring
Monitoring A fix has been deployed and {affected surface} is recovering. We're monitoring closely to confirm full recovery before we resolve this incident. If you're still seeing issues, please contact {support channel}. Next update by {time UTC}.
Resolved
Resolved This incident is resolved and {affected surface} has been operating normally since {time UTC}. The impact window was {start}–{end} UTC, during which {plain summary of impact}. {If applicable: "No customer action is required." / "If you {action needed}…"} We're sorry for the disruption. A post-mortem will follow {timeframe} — thank you for your patience.
Degraded performance (partial)
Investigating — {Surface} slower than normal {Surface} is responding slower than usual for {scope}. Requests are completing, but you may see {symptom, e.g. "timeouts on large exports"}. We're investigating. Next update by {time UTC}.
Scheduled maintenance
Scheduled maintenance — {date}, {start}–{end} UTC We'll be performing maintenance on {surface}. Expected impact: {none / brief interruptions of up to X minutes}. We'll confirm here when maintenance is complete.
Internal vs public comms
Run two channels from the first minute, and never let them blur:
- Internal (war-room chat, incident doc): raw hypotheses, real service names, dashboards, blunt speculation, timestamps for the post-mortem. Fast and messy is correct here.
- Public (status page, subscriber emails): only confirmed facts, customer language, stage + cadence. One person — a comms lead who is not the person debugging — translates from the first channel to the second.
The single rule that prevents most public-comms disasters: nothing speculative crosses the wall. Internal says "maybe the failover corrupted replication"; public says "we've identified a database issue" only once it's confirmed.
Wire it up before you need it
Templates only help if the incident reaches you before your customers' tweets do. CompleteStatus detects downtime with confirmation probes, opens the incident, and posts to your public status page with the Investigating → Resolved workflow built in — with subscriber notifications, multi-channel alerting (Slack, Discord, Telegram, email, webhooks), and uptime, SSL, DNS and security monitoring feeding the same incident timeline instead of five separate tools. Check any endpoint right now with the free uptime test, then set up monitoring and a status page free — the free tier allows commercial use.