Skip to main content

The Fastly Outage — What Tuesday's Hour of Broken Internet Says About Third-Party Risk

Fastly's June 8 outage took down Amazon, Reddit and gov.uk for nearly an hour. What one config change teaches about CDN blast radius and third-party risk.

The CompleteStatus Team 6 min read

The Fastly Outage — What Tuesday's Hour of Broken Internet Says About Third-Party Risk

Want this checked continuously?

CompleteStatus grades your headers, SSL and email security 24/7 — free, commercial use allowed.

Monitor your site free

On Tuesday morning, a sizable chunk of the internet simply stopped. Starting just before 10:00 UTC on June 8, visitors to Amazon, Reddit, Twitch, Spotify, PayPal's site, Stack Overflow, the New York Times, the Guardian, CNN and the UK government's gov.uk, among many others, got bare 503 Service Unavailable errors instead of pages. The cause wasn't a cyberattack or a fiber cut. It was Fastly, the CDN sitting in front of all of them.

Fastly's summary, published the same day, is admirably direct: a software deployment on May 12 introduced a latent bug, and on June 8 a single customer pushed a perfectly valid configuration change that happened to trigger it, causing roughly 85% of Fastly's network to return errors. Detection reportedly took about a minute, the fix was rolling out within three quarters of an hour, and most of the network was serving normally again inside the hour. As global outages go, the response was fast. Still, for most of an hour, one customer's config change turned off a measurable fraction of the web. Two days on, here's what the rest of us should take from it.

What actually happened

The sequence is worth knowing because of how ordinary every step was:

  • May 12: Fastly ships a software update containing an undiscovered bug. Nothing happens for four weeks; every test and every customer interaction passes.
  • June 8, ~09:47 UTC: a customer pushes a valid configuration change that hits the bug's exact trigger conditions. The failure cascades across the network.
  • ~10:27–11:00 UTC: Fastly identifies the trigger, disables the offending configuration, and the network comes back as caches refill.

No attacker, no negligence, no "human error" in the usual sense. A latent bug plus a legitimate input: the kind of failure that no amount of vendor diligence can promise you'll never see again, at Fastly or anyone else.

Your blast radius is bigger than your architecture diagram

The uncomfortable exercise this week is inventory. Most production sites in 2021 have quietly accumulated a stack of single points of failure that don't appear on any diagram because they're someone else's infrastructure:

  • CDN in front of everything. When it's down, it doesn't matter how healthy your origin is.
  • DNS provider. If resolution fails, nothing else gets a chance to.
  • TLS termination, WAF, DDoS protection, often the same vendor as the CDN, concentrating the risk further.
  • Third-party JavaScript: payment widgets, chat, analytics, fonts served off someone's edge. Tuesday's outage reportedly broke parts of many sites that didn't front with Fastly at all, because a critical script did.

For each dependency, ask two questions: what exactly breaks for my users when this vendor has a bad hour? and how would I even know it's them and not me? If the answer to the second is "I'd check Twitter," read on.

Monitor from outside your own infrastructure

Plenty of teams affected on Tuesday likely had monitoring that stayed green the whole time, because it was watching the origin. Health checks inside your own network, or synthetic checks pointed at the origin's direct address, measure a path your users never take. Their requests go through DNS, the CDN edge and the third-party stack, and on Tuesday that's where the failure lived.

The principle: monitor the path your users travel, from where your users are. That means external checks against your public, CDN-fronted URLs from multiple geographic locations, catching regional edge failures a single vantage point misses. Keep a check on the origin too, and the combination becomes a diagnosis: edge failing while origin answers is the "it's the CDN" signal. That turns an hour of confusion into an immediate, correct response, starting with not restarting servers that were never broken.

Your status page can't live in your blast radius

Tuesday had the usual irony: some affected companies couldn't tell their users what was happening because their status page rode the same infrastructure as their product.

The rule is old but apparently needs restating every time this happens: the thing that reports your downtime must not share your downtime. Use different hosting, ideally a different CDN or none, and think hard about DNS: a status page on a subdomain of your main zone dies with your DNS provider, however carefully you separated the hosting. During an incident, a reachable status page is the difference between "they know, they're on it" and a support inbox melting down, and it buys your team the quiet to actually fix things.

Have a bypass plan you've actually rehearsed

You can't make a third party infallible, but you can decide in advance what you'll do during their bad hour:

  • Know how to take the CDN out of the path. For many setups it's a DNS change repointing traffic at the origin (or a secondary provider). Write down the steps, note which TTLs must already be low for this to work in minutes rather than hours, and mind the details: origin certs valid for direct traffic, firewall rules that currently only admit CDN IPs.
  • Accept the degraded mode. Your origin probably can't take full production load unprotected. That's fine. Slow beats down, and a static "we're degraded" page you can serve from anywhere beats both.
  • Rehearse it once. A bypass plan first attempted during the outage is a hope, not a plan.
  • Decide your threshold. Switching costs something (cache misses, load, TLS details). Knowing in advance that "we bypass after X minutes of confirmed edge failure" beats debating it live.

Trust vendors; verify independently

None of this is an argument against CDNs. Fastly's network absorbs attacks and delivers performance no single-origin setup can match, and their transparency this week deserves the credit it's getting. It's an argument for the old rule: trust, but verify. Every third party in front of your site is part of your uptime, and your visibility into it shouldn't depend on the vendor's own dashboard, which may be having a bad hour too.

CompleteStatus checks your site from outside, along the full user path through your CDN and DNS, so on the next bad Tuesday you know within a minute what's broken and whose it is. Those bare 503s the world stared at this week are exactly the kind of signal worth reading correctly (we wrote about status codes through a monitoring lens). Create a free account and get an independent view of everything you depend on, including the parts you don't run.

Written with AI assistance and reviewed by the CompleteStatus team.