Skip to main content

AWS in October, Cloudflare in November: Two Outages, One Lesson About Hidden Dependencies

A DynamoDB DNS race in us-east-1 and an oversized Cloudflare config file took down much of the web within a month. What both teach about hidden dependencies.

The CompleteStatus Team 6 min read

AWS in October, Cloudflare in November: Two Outages, One Lesson About Hidden Dependencies

Want this checked continuously?

CompleteStatus grades your headers, SSL and email security 24/7 — free, commercial use allowed.

Monitor your site free

In the space of a month, two of the most depended-upon pieces of internet infrastructure had bad days. On October 20, a failure in AWS's us-east-1 region knocked out a long list of apps and services for most of a working day. On November 18, Cloudflare's network started returning errors for a large share of the traffic it proxies, taking a lot of well-known sites down with it for several hours.

Both companies have now published detailed write-ups, and both are worth reading in full. This post isn't a retelling. It's about what the two incidents have in common, and what a team that isn't AWS or Cloudflare can actually do with that.

What happened at AWS

According to AWS's post-event summary, the root cause was a latent race condition in the automation that manages DNS records for DynamoDB in us-east-1. Two parts of that system, operating on different versions of a DNS plan, interacted in a way that left the regional DynamoDB endpoint's DNS record empty. Anything that looked up that name, including AWS's own internal services, couldn't find DynamoDB.

The DNS record was repaired within a few hours. The outage lasted much longer, because other AWS systems depend on DynamoDB and had got into bad states while it was unreachable. AWS described knock-on problems launching new EC2 instances and with Network Load Balancer health checks, which took most of the rest of the day to fully work through. Services that had nothing obviously to do with DynamoDB were affected because AWS's own control plane depends on it.

What happened at Cloudflare

Cloudflare's write-up, published the same day, traced its outage to a change in database permissions that caused a query to return duplicate rows. That query generates a configuration file used by Cloudflare's Bot Management system. The file roughly doubled in size, exceeded a limit hard-coded in the software that reads it, and the proxy handling customer traffic started failing. Because the file was regenerated every few minutes, and only some database nodes had the new permissions at first, the network flapped between working and broken for a while, which initially looked more like an attack than a bug.

Once the cause was found, Cloudflare stopped generating the bad file, pushed a known-good version and restarted the affected services. Most traffic recovered within a few hours of the start, with a longer tail after that.

The shared pattern

Two different companies, two different bugs. The shape is the same:

  • A routine, automated change (DNS plan updates, a permissions change) that had presumably happened many times before without incident.
  • Bad data propagated at machine speed into a system that trusted it: an empty DNS record, an oversized config file.
  • A blast radius much larger than the component that failed, because other systems depended on it in ways that weren't obvious from the outside.

That third point is the one that matters for everyone else. Plenty of teams that were hit on October 20 would have told you, the week before, that they didn't depend on DynamoDB. Many of them didn't, directly. They depended on a SaaS vendor, a login provider, a payment processor or a feature-flag service that ran in us-east-1. On November 18, many sites that went down weren't "on Cloudflare" in any sense their own engineers thought about day to day. A widget, an auth flow or an API they call was.

We've made versions of this point before, after the S3 outage in 2017, the Dyn DDoS in 2016 and CrowdStrike last year. It keeps being true because dependency graphs keep getting deeper faster than anyone maps them.

What you can do that doesn't require a second cloud

"Go multi-region" or "go multi-CDN" is the standard advice after an outage like these, and for most small and mid-sized teams it's the wrong first move. It's expensive, it adds its own failure modes, and it doesn't help when the thing that broke is a vendor you don't control. Cheaper steps come first.

1. Write down your dependencies, including the indirect ones

Make a list of every external service your product needs to serve a request or complete a sign-up, checkout or login. Then, for each one, find out where it runs. Vendor status pages and trust pages often say which cloud and region they use. You won't get a complete picture, but you'll find the surprises: three "independent" vendors that all sit in the same region, for example.

2. Monitor the paths users take, not just your origin

If your origin was healthy on October 20 but your login provider wasn't, your users couldn't log in and a check against your homepage stayed green. Checks that exercise the real flows (log in, load the dashboard, hit the API endpoint that calls the payment provider) catch this class of failure. So does a health check endpoint that reports on critical dependencies separately from the app itself, as long as you're careful not to make your whole site "down" because an optional dependency is.

3. Know what degraded mode looks like

For each critical dependency, decide in advance what happens when it's gone. Can you show cached content? Queue orders and process them later? Hide the recommendation widget instead of failing the whole page? The cheapest resilience is usually a timeout and a fallback in the right place, not a second provider.

4. Keep your status communication off the thing that failed

If your status page, your incident chat and your alerting all run on the same provider as your product, a provider outage takes out your ability to talk about the outage. Check where each of those actually lives.

5. Watch the watchers

Your monitoring is itself a dependency. If it runs in the same region as your production stack, a regional failure can blind you at exactly the wrong moment. You don't need to fix everything, just know which outages you would and wouldn't see.

Timing

It's Thanksgiving week in the US, and many teams are in a change freeze until after Cyber Monday. That makes this a bad week for infrastructure changes and a good week for the parts of this list that are just reading and writing things down. If you worked through our peak season checklist earlier this month, the dependency inventory is the natural next step.

Neither of these outages was caused by anything their customers did, and neither could have been prevented by them. What customers controlled was how long it took to notice, how clearly they communicated, and whether a vendor's bad day took down the whole product or just one feature. That's still worth working on before the next one.

Written with AI assistance and reviewed by the CompleteStatus team.