Articles
AWS in October, Cloudflare in November: Two Outages, One Lesson About Hidden Dependencies
A DynamoDB DNS race in us-east-1 and an oversized Cloudflare config file took down much of the web within a month. What both teach about hidden dependencies.
2024: The Year in Outages, and What It Taught Us
CrowdStrike, polyfill.io and a year of cloud and carrier failures: what 2024's big outages share, and what small teams should carry into 2025.
The CrowdStrike Outage — Lessons From the Biggest IT Failure Ever
A bad CrowdStrike channel file blue-screened ~8.5 million Windows machines on July 19. What may be the largest IT outage ever teaches teams of every size.
The Rogers Outage: When One Carrier Takes Down Phones, Payments and 911
Friday's Rogers outage knocked out mobile, internet and Interac debit across Canada. What it says about single-carrier risk and where your alerts go.
The Atlassian Outage — Two Weeks of Downtime and What Not to Say
A maintenance script permanently deleted ~400 Atlassian customers' cloud sites; restoration took two weeks. Lessons on incident communication and backups.
The Facebook Outage — What BGP and DNS Taught Everyone This Week
Facebook, Instagram and WhatsApp vanished for six hours this week after a BGP misstep took out their DNS. What the outage teaches every team about resilience.
The Fastly Outage — What Tuesday's Hour of Broken Internet Says About Third-Party Risk
Fastly's June 8 outage took down Amazon, Reddit and gov.uk for nearly an hour. What one config change teaches about CDN blast radius and third-party risk.
One Bad Regex: Lessons From the July 2 Cloudflare Outage
A WAF rule with a backtracking regex pushed Cloudflare's edge CPUs to 100% and served 502s worldwide. What the outage teaches every team.
Yesterday's BGP Route Leak: How a Small ISP Rerouted the Internet
A Pennsylvania ISP's route optimizer leaked routes, Verizon propagated them, and traffic for Cloudflare and Amazon went sideways. What BGP teaches us.
Sunday's Google Cloud Outage: When Your Provider's Network Has a Bad Afternoon
A Google configuration change on June 2 congested networking across US regions for hours. What it teaches about shared fate, alert paths and outside-in checks.
GitHub's Day-Long Outage: Choosing Consistency Over Availability
GitHub spent about a day degraded this weekend and chose data integrity over a fast recovery. What that trade-off means, and what to check in your own pipeline.
1.35 Terabits Per Second: What the GitHub DDoS Teaches Everyone Else
How a 1.35 Tbps memcached amplification attack knocked GitHub offline for about ten minutes, and what the record DDoS teaches every team about preparation.
More topics
Get notified when CompleteStatus opens
New accounts are closed while we're in private beta. Leave your email and we'll send one message the moment sign-ups open — nothing else.