Skip to main content

Incident Communication Templates You Can Copy-Paste at 3 A.M.

Copy-paste incident updates for every stage (investigating, identified, monitoring, resolved), plus the tone rules that keep customer trust.

The CompleteStatus Team 5 min read

Incident Communication Templates You Can Copy-Paste at 3 A.M.

Want this checked continuously?

CompleteStatus grades your headers, SSL and email security 24/7 — free, commercial use allowed.

Monitor your site free

Good incident communication is written before the incident. At 3 a.m., with the database down and three people on a call, nobody produces calm, precise customer-facing prose from a blank page. They either go silent, which is the worst option, or write something panicked they'll regret. The fix is boring and it works: templates for every stage, agreed on in daylight, filled in under pressure.

Below is a full set (first update, progress, resolution, follow-up), the tone rules that make them work, and the internal/public split that keeps the war room honest.

The first update: acknowledge within minutes

The first update has one job: end the silence. Customers hitting errors want to know whether it's just them and whether you know. Every minute without an answer turns into support tickets and doubt about renewing. Aim to post within 5–10 minutes of confirming impact, on your status page first and everywhere else second.

Include:

  • That you know, and you're on it
  • What's affected, in customer language ("API requests," "dashboard logins," not internal service names)
  • Rough scope if you know it ("some users," "all requests to…")
  • When the next update is coming, and then keep that promise

Omit:

  • Root-cause speculation ("we believe a database failover…"). Early guesses are often wrong, and correcting them costs trust
  • ETAs you can't stand behind
  • Blame, whether of a vendor, a deploy or a person

A vague acknowledgement posted now beats a precise one posted late.

Progress updates: cadence is the message

Once the first post is up, the cadence is the communication. State it explicitly ("next update by 14:30 UTC") and hit it even when nothing has changed. "Still investigating, no change" is information; silence reads as abandonment. As a default, update every 30 minutes for a full outage and every 60 for partial degradation. Move through explicit stages so progress is easy to scan: Investigating → Identified → Monitoring → Resolved.

Resolution and follow-up

Resolve only when customers have recovered, not when the fix is deployed; that's what the Monitoring stage is for. The resolution note should explain what happened in about a paragraph, give the impact window in UTC, and say whether any customer action is needed.

For significant incidents, follow up within a few days with a short public post-mortem: timeline, root cause in plain language, and what you're changing. Done honestly, these build trust rather than damage it. The companies known for handling incidents well are known for it largely because of their write-ups.

Tone rules

  1. No blame. Not "our hosting provider failed"; customers bought from you. "A component of our infrastructure" is honest without pointing fingers. Internally, blamelessness is what keeps engineers reporting mistakes early.
  2. No over-promising. Never "resolved shortly," and no ETA you're less than 90% sure of. The only safe promise is the next update time.
  3. Plain language. "Logins are failing," not "elevated error rates on the authentication layer."
  4. Calm, not casual. No exclamation marks, no "oops!", no memes. Warmth belongs in the apology, not the diagnosis.
  5. Own it. "We're sorry, this is on us" reads better in every follow-up than a passive "service interruptions were experienced."

Copy-paste templates

Swap the {placeholders}, keep the structure.

Investigating (the first post)

Investigating: {Short, plain title, e.g. "API requests failing"} We're aware that {affected surface} is {failing / degraded} for {scope, e.g. "some users" / "all requests"} and are actively investigating. {If true: "Your data is not affected."} Next update by {time UTC}.

Identified

Identified We've identified the cause: {one plain-language sentence, no blame, e.g. "a configuration change affecting our database connections"}. A fix is {being prepared / rolling out now}. {Affected surface} remains {status}. Next update by {time UTC}.

Monitoring

Monitoring A fix has been deployed and {affected surface} is recovering. We're monitoring closely to confirm full recovery before we resolve this incident. If you're still seeing issues, please contact {support channel}. Next update by {time UTC}.

Resolved

Resolved This incident is resolved and {affected surface} has been operating normally since {time UTC}. The impact window was {start}–{end} UTC, during which {plain summary of impact}. {If applicable: "No customer action is required." / "If you {action needed}…"} We're sorry for the disruption. A post-mortem will follow {timeframe}. Thank you for your patience.

Degraded performance (partial)

Investigating: {Surface} slower than normal {Surface} is responding slower than usual for {scope}. Requests are completing, but you may see {symptom, e.g. "timeouts on large exports"}. We're investigating. Next update by {time UTC}.

Scheduled maintenance

Scheduled maintenance: {date}, {start}–{end} UTC We'll be performing maintenance on {surface}. Expected impact: {none / brief interruptions of up to X minutes}. We'll confirm here when maintenance is complete.

Internal vs public comms

Run two channels from the first minute and never let them blur:

  • Internal (war-room chat, incident doc): raw hypotheses, real service names, dashboards, blunt speculation, timestamps for the post-mortem. Fast and messy is correct here.
  • Public (status page, subscriber emails): only confirmed facts, customer language, stage and cadence. One person, a comms lead who is not the person debugging, translates from the first channel to the second.

The rule that prevents most public-comms disasters: nothing speculative crosses the wall. Internal says "maybe the failover corrupted replication"; public says "we've identified a database issue," and only once it's confirmed.

Wire it up before you need it

Templates only help if you hear about the incident before your customers start posting about it. CompleteStatus marks a monitor down after two consecutive failed checks, opens an incident with a timeline, alerts your channels (Slack, Discord, Telegram, email, webhooks and more), and shows the outage on your public status page, where visitors can subscribe to email updates. Uptime, SSL, DNS and security monitoring all feed the same incident list. The update text is still yours to write, which is what the templates above are for. Check any endpoint now with the free uptime test, or start on the free plan, which allows commercial use.

Written with AI assistance and reviewed by the CompleteStatus team.