Delta and Southwest: When the Backup Data Center Isn't a Backup
Two airlines, two data center failures, thousands of cancelled flights in a month. What they suggest about failover plans that exist mostly on paper.
Want this checked continuously?
CompleteStatus grades your headers, SSL and email security 24/7 — free, commercial use allowed.
In the space of three weeks, two of the largest airlines in the United States grounded themselves with their own IT.
On July 20th, Southwest Airlines had a technology failure that the company has attributed to a faulty network router. The knock-on effects ran for days, and more than two thousand flights were reported cancelled before the schedule recovered. Then, early on the morning of August 8th, Delta lost power at its data center in Atlanta. Delta has said a critical power control module failed, and that when it did, critical systems and network equipment didn't switch over to backups as they were supposed to. Check-in, boarding and crew scheduling went down worldwide. Delta cancelled roughly two thousand flights over the following days, according to its own updates, and was still apologising to customers at the end of the week.
We don't have postmortems for either event, and we may never get detailed public ones. Airlines aren't obliged to publish them. What we do have is the shape of the failures, and the shape is familiar to anyone who has run production systems for long enough: a single component fails, a redundancy plan exists, and the redundancy plan doesn't do what everyone assumed.
"We have a backup" is a claim, not a fact
Nearly every organisation of Delta's size has a disaster recovery plan. It will name a secondary site, a set of systems that fail over, and a target recovery time. It will usually have been reviewed and signed off.
What the plan can't tell you is whether it still works. Infrastructure drifts constantly. Servers are added in a hurry and plugged into whatever power strip has a free socket. A network change moves a critical service onto a path that doesn't exist at the secondary site. A dependency nobody documented, a licence server or an internal DNS resolver, lives only in the primary building. None of this shows up until the day the primary goes away.
The only reliable way to know a failover works is to fail over. Deliberately, on a schedule, while people are awake and watching. Organisations that do this find problems every time they try, which is the point. Organisations that don't find the same problems during a real outage, with every customer watching.
For smaller teams the equivalent questions are simpler but just as uncomfortable:
- If your primary database server died now, how long until the replica is serving writes? Has anyone actually promoted it, or only read the documentation for how to?
- If your hosting provider's region went dark, is there anywhere else your site can run? If the honest answer is "no, we'd wait", that's a legitimate choice, but it should be a known choice.
- Does anything critical depend on one physical thing: one router, one power feed, one disk, one person's laptop?
Recovery is the long part
What stands out in both airline outages is how long the tail was. The technical failure at Delta was reportedly resolved within hours. The operation took days to recover.
That's typical of systems with a lot of state. When a scheduling system goes down, aircraft and crews end up in the wrong places. When it comes back, it has to reconcile what it thinks happened with what actually happened. Queued work floods in all at once. Staff who were improvising with paper have to be brought back onto the system. Each of those steps takes time that no failover plan accounts for, because failover plans are written about the moment of failure, not the week after it.
Web systems have smaller versions of the same tail. A payment backlog that replays when the processor comes back. A job queue that has piled up for four hours and now saturates the database the moment workers restart. Caches that were cold when traffic returned all at once. If your plan ends at "service restored", it's missing the part where things are most likely to break a second time.
Measure from outside, because inside will be down too
One detail that gets overlooked in outages like these: when the data center goes dark, so does most of the tooling you'd use to understand it. Dashboards hosted in the same building disappear. Alerting that runs on the same network can't send anything. The first people to know are often customers, and then the press.
We've written before about why green dashboards lie: internal metrics describe the parts, not the service your users actually get. The airline outages add a second reason to measure from outside. External checks don't share a failure domain with the thing they're checking. If your whole primary site loses power, a check running somewhere else still notices, still records when it started, and still tells someone.
That record matters later too. "When exactly did it start, and when was it really back for users?" is the first question in every post-incident review, and the one internal logs answer worst, because the logging stopped when the power did.
Practical things to do this month
You don't need an airline's budget to take something useful from this. A short list:
- Write down your single points of failure. Not the ones you've fixed, the ones you haven't. For each, note what happens if it fails and how long recovery would take. Share the list with whoever makes budget decisions. Some of these you'll accept. That's fine, as long as it's on purpose.
- Test one failover. Pick the one you're least sure about. Promote the replica on a staging copy, restore last night's backup to a scratch server, or switch DNS to the standby and see what breaks. Time it. Write down every surprise.
- Check the power and network assumptions. If you run your own hardware, confirm that every server with a redundant power supply actually has both supplies connected, to different feeds. It's a boring walk through a rack and it's exactly the kind of thing that drifts.
- Plan for the tail. For each critical queue or backlog, decide what happens when it comes back with hours of work in it. Rate limits on replay, a way to drain slowly, and a plan for manual cleanup.
- Make sure someone outside your infrastructure is watching. At minimum, one check of your public site from a location that doesn't share your network or power, alerting to a channel that doesn't depend on your own mail server.
- Decide who talks to customers. Delta and Southwest both spent days communicating about their outages. For a web business the equivalent is a status page and a clear owner for updates. Decide who that is before you need them.
Redundancy you haven't exercised is a guess
The airline outages will be studied for a while, and the details will eventually be more interesting than what's public today. The general lesson doesn't need the details: redundancy is a property you demonstrate, not one you design. Every untested failover path is an assumption, and assumptions age badly.
External monitoring won't fix an untested failover. What it will do is tell you, within a minute or two and from outside the blast radius, that one has just failed, so the recovery starts from a known time instead of from the first angry customer. That part, at least, is something we can help with at CompleteStatus.
Written with AI assistance and reviewed by the CompleteStatus team.
Get notified when CompleteStatus opens
New accounts are closed while we're in private beta. Leave your email and we'll send one message the moment sign-ups open — nothing else.