Skip to main content

Incidents guides and articles

Every CompleteStatus blog post tagged Incidents — 25 articles, newest first.

Articles

outages dns

The Facebook Outage — What BGP and DNS Taught Everyone This Week

Facebook, Instagram and WhatsApp vanished for six hours this week after a BGP misstep took out their DNS. What the outage teaches every team about resilience.

6 min read
outages cdn

The Fastly Outage — What Tuesday's Hour of Broken Internet Says About Third-Party Risk

Fastly's June 8 outage took down Amazon, Reddit and gov.uk for nearly an hour. What one config change teaches about CDN blast radius and third-party risk.

6 min read
backups disaster-recovery

The OVHcloud Strasbourg Fire: Were Your Backups in the Same Building?

Last week's fire destroyed a data centre in Strasbourg and took backups with it. How to check where your copies really live, and how fast you could restore.

6 min read
security supply-chain

SolarWinds Orion: When the Monitoring Tool Is the Way In

A trojanized Orion update reached thousands of networks. What we know so far, and what it means for any tool that holds the keys to your infrastructure.

6 min read
cloudflare outages

One Bad Regex: Lessons From the July 2 Cloudflare Outage

A WAF rule with a backtracking regex pushed Cloudflare's edge CPUs to 100% and served 502s worldwide. What the outage teaches every team.

6 min read
bgp outages

Yesterday's BGP Route Leak: How a Small ISP Rerouted the Internet

A Pennsylvania ISP's route optimizer leaked routes, Verizon propagated them, and traffic for Cloudflare and Amazon went sideways. What BGP teaches us.

6 min read
outages cloud

Sunday's Google Cloud Outage: When Your Provider's Network Has a Bad Afternoon

A Google configuration change on June 2 congested networking across US regions for hours. What it teaches about shared fate, alert paths and outside-in checks.

7 min read
outages github

GitHub's Day-Long Outage: Choosing Consistency Over Availability

GitHub spent about a day degraded this weekend and chose data integrity over a fast recovery. What that trade-off means, and what to check in your own pipeline.

6 min read
security ransomware

WannaCry: The Patch Existed for Two Months. Why Wasn't It Installed?

WannaCry is spreading through networks using a flaw Microsoft patched in March. What the worm teaches about patch cadence, EOL systems and tested backups.

5 min read
aws outages

The S3 Outage: Why Your Status Page Can't Live on Your Own Infrastructure

Amazon's four-hour S3 outage broke a large part of the web, and AWS's own dashboard couldn't turn red. Why status pages need independent infrastructure.

6 min read
security cloudflare

Cloudbleed: What a Leaking CDN Teaches About Shared Infrastructure

Cloudflare's edge leaked memory from unrelated sites into served pages for months. What Cloudbleed teaches about shared infrastructure and secrets.

6 min read
backups databases

The GitLab.com Database Incident: A Backup You Haven't Restored Is Just a Hope

Last week GitLab.com lost hours of production data and found its backups weren't working. What happened, and how to check your own before you need them.

7 min read

More topics