Blog / The Atlassian Outage — Two Weeks of Downtime and a Masterclass in What Not to Say
outages incidents communication backups

The Atlassian Outage — Two Weeks of Downtime and a Masterclass in What Not to Say

A maintenance script permanently deleted ~400 Atlassian customers' cloud sites; restoration took two weeks. Lessons on incident communication and backups.
The CompleteStatus Team · · 6 min read
The Atlassian Outage — Two Weeks of Downtime and a Masterclass in What Not to Say
Want this checked continuously?
CompleteStatus grades your headers, SSL and email security 24/7 — free, commercial use allowed.
Monitor your site free

For the past two weeks, roughly 400 companies have been living a very specific nightmare: their Jira, Confluence and other Atlassian cloud products didn't go down — they were deleted. In early April, a maintenance script intended to deactivate a deprecated standalone app was fed the wrong identifiers and, instead of removing the app, permanently deleted the entire cloud sites of about 400 customers. Restoration has taken up to two weeks, and only now — as we write this — are the last affected customers reportedly coming back online.

The engineering failure is instructive. The communication failure is the part everyone will remember. Both carry lessons for teams a thousandth of Atlassian's size — arguably especially for those teams.

What happened, as best we know

Atlassian has acknowledged the outline publicly: a script meant to delete a deprecated app was run with a mode and a target list that didn't match — it received IDs for entire cloud sites rather than the app, and it ran in permanent-deletion mode rather than the recoverable soft-delete. The data wasn't lost forever — Atlassian keeps backups — but restoring several hundred customers' sites individually, without rolling back everyone else's data on the same shared infrastructure, turned out to be a slow, largely manual process. Hence the astonishing headline: a top-tier SaaS vendor quoting restoration times of up to two weeks.

Meanwhile, affected customers spent the first several days with something worse than bad news: almost no news. Early communication was widely criticized as slow, generic and impersonal — support tickets that couldn't be filed because the ticketing product was the thing that was down, days of near-silence for some customers, and status updates that leaned on the phrasing that this affected "only" a small percentage of customers.

Lesson 1: "a small percentage of customers" is a phrase to ban

Statistically, Atlassian was right — around 400 sites out of hundreds of thousands is well under 1%. Rhetorically, it was a disaster, because of an asymmetry every incident communicator needs tattooed somewhere visible: for each affected customer, the outage is 100%. A company whose entire project history, documentation and ticket queue has vanished does not experience your incident as a rounding error — and being told, implicitly, "this barely matters in aggregate" converts anxiety into anger.

The fix costs nothing: describe impact from the affected user's perspective ("if your site is affected, it is fully unavailable; here is what we know and what to expect"), and save the percentage for the postmortem, where it belongs as scope, not as reassurance.

Lesson 2: communication cadence is a commitment, not a mood

The deepest customer complaint these two weeks hasn't been the deletion — accidents happen — it's the silence. Days without a substantive update, and no personalized contact for organizations facing an existential data question, reads as indifference even when hundreds of engineers are working around the clock behind the scenes. Which, by all accounts, they were. Effort that isn't communicated might as well not exist.

The mechanics that separate good incident comms from this:

  • Commit to a cadence and state it. "Next update by 16:00 UTC" — and then update at 16:00 UTC even if the update is "no material change, still restoring, next update at 20:00." A promised update that arrives with nothing new builds more trust than an unpromised one that arrives with news.
  • Have an out-of-band contact path. When your product is the support channel, its outage takes your comms down too — the same fate-sharing trap as hosting your status page on your own infrastructure. Email lists, a status page on separate infrastructure, and phone contacts for your largest accounts need to exist before the incident.
  • Give honest time estimates, even when they're ugly. Atlassian eventually said "up to two weeks," and it was received far better than the vague reassurances that preceded it — because customers can plan around an ugly truth. They can't plan around "soon." Under-promise, over-deliver, and never let legal polish strip the specifics out of an engineering update.
  • Sign it like a human. Updates that read like they cleared a communications-review committee land worse than two plain sentences from a named engineer. Speed and sincerity beat polish in an incident, every time.

Lesson 3: backups you haven't restored are a hypothesis

The uncomfortable engineering lesson: Atlassian had backups — this was never a data-loss story — and restoration still took two weeks, because restoring a subset of tenants out of shared, multi-tenant infrastructure had apparently never been exercised at this scale or shape. The backup existed; the restore path for this scenario didn't, operationally speaking.

For your own systems, the checklist writes itself:

  • Test restores, not backups. A backup job that runs green nightly proves nothing about your ability to get data back. Schedule actual restore drills — to a scratch environment, timed — and treat the measured duration as your real RTO, not the one in the disaster-recovery doc.
  • Rehearse partial restores. Full-system recovery is the easy case. "Restore one customer / one table / one folder to how it looked Tuesday, without touching anything else" is the request you'll actually get, and it's much harder.
  • Make destructive operations awkward. The trigger here was a script fed the wrong IDs in the wrong mode. Soft-delete with a retention window as the default, permanent deletion as a separate privileged path, dry-run output reviewed by a second person for anything batch-destructive — friction on the irreversible path is a feature.

Lesson 4: keep your own record of vendor downtime

A quieter lesson for everyone who depends on SaaS vendors: during those two weeks, an affected customer's ability to answer basic questions — when did it start for us? how many hours have we lost? does this breach the SLA we're paying for? — depended largely on the vendor's own account of events. Vendor status pages are written by the party with the least incentive to make the outage look long, and they famously round down.

Independent monitoring of the services you rely on gives you your own timestamped record: when the outage started for you, when it actually ended for you, and how the vendor's narrative compares. That's leverage in SLA conversations, evidence for your own customers when a vendor outage cascades into yours, and — as we wrote after October's Facebook outage — the only view of availability that matches what users experience.

Watch what you depend on

CompleteStatus monitors any URL you can reach — including your vendors' endpoints and your own product's dependencies — with uptime history you own, independent of anyone's status page. Alerts land in email, Slack or webhooks the moment a check fails, and the timeline is yours to export when the SLA conversation happens. Create a free account and add a monitor for the one SaaS product your company genuinely can't work without. This month, of all months, you can probably name it instantly.

Stop checking by hand

CompleteStatus runs these exact checks around the clock — uptime, SSL, DNS, security headers and SPF/DKIM/DMARC — and alerts you the moment something changes. One dashboard, one bill.
Start free Run a free check
Uptime is table stakes. We watch the rest — security headers, email authentication, certs and DNS, with the fix attached.
Start free
Company
© 2026 CompleteStatus. All rights reserved. CompleteStatus — operated in the United States · support@completestatus.com