Skip to main content

The S3 Outage: Why Your Status Page Can't Live on Your Own Infrastructure

Amazon's four-hour S3 outage broke a large part of the web, and AWS's own dashboard couldn't turn red. Why status pages need independent infrastructure.

The CompleteStatus Team 6 min read

The S3 Outage: Why Your Status Page Can't Live on Your Own Infrastructure

Want this checked continuously?

CompleteStatus grades your headers, SSL and email security 24/7 — free, commercial use allowed.

Monitor your site free

Last Tuesday, Amazon S3 in the us-east-1 region went down for roughly four hours, and we got a live demonstration of how much of the web is built on one storage service in northern Virginia. Images vanished from major sites, file uploads failed, smart-home devices stopped responding, and applications whose owners would have said they "don't really use AWS" stopped working.

The detail most people will remember: for a good part of the outage, AWS's own status dashboard showed green. Nobody was hiding anything. The red and yellow status icons were themselves hosted on S3, so the status page couldn't report the outage it was part of. AWS ended up posting updates on Twitter.

What happened, in Amazon's words

Amazon's postmortem is admirably direct. An engineer on the S3 team, debugging a slowdown in the S3 billing system, ran a command from an established playbook meant to remove a small number of servers from one subsystem. One input was entered incorrectly, and the command removed a much larger set of servers than intended, including servers supporting S3's index and placement subsystems. Those are the systems that know where every object lives and where new ones should go.

Both subsystems needed a full restart. S3 in us-east-1 had grown enormously over the years without those systems being fully restarted, and the restart and safety checks took hours rather than minutes.

There's no villain here, no attack and no exotic hardware failure. It's a typo, made with legitimate credentials, inside a documented procedure. Amazon's fixes (tooling that limits how quickly capacity can be removed, and work to make those subsystems recover faster) are process fixes, because it was a process failure. If your own runbooks let one mistyped argument remove unbounded capacity, the same thing can happen to you.

The dashboard that couldn't turn red

A status page has one job: keep working when everything else doesn't. That means it should share as little fate with production as you can arrange:

  • Different infrastructure, ideally a different provider. If your app runs on AWS, a status page on S3, or on anything in the same region, will fail alongside it. Host it somewhere whose bad day is unlikely to be the same day.
  • Different DNS. If your nameservers are down, status.yourapp.com on the same nameservers is down too. A separate DNS provider, or at least a well-cached, long-TTL record, keeps it reachable.
  • No dependencies on your stack. The status page shouldn't need your database, API, asset pipeline or single sign-on to render. Static and self-contained is right.
  • Updatable from a phone on coffee-shop Wi-Fi. During a serious incident your VPN, deploy pipeline or office network may be part of the problem. If posting an update requires production access, it won't get posted.

The test is simple: imagine production is gone and ask whether customers can still find out what's going on. Amazon, with more infrastructure than almost anyone, failed that test in public last week. It takes five minutes to check whether you'd pass.

Do you know what depends on S3?

The second surprise of the day was how many teams found dependencies they didn't know they had. The app was "up" and half its features weren't. It's worth literally grepping for them:

grep -rn --include="*.php" --include="*.js" --include="*.yml" \
  -e "s3.amazonaws.com" -e "s3-us-" -e "AWS_BUCKET" .

The usual suspects from last Tuesday:

  • User uploads and avatars. Pages render, every image is broken, every upload form errors.
  • Static assets and downloads. JS bundles, PDFs, installers, podcast audio served from a bucket.
  • Background jobs. Exports, imports, report generation and backups that read or write S3 and failed quietly.
  • Third parties. Your support widget, payment provider or analytics script may rely on S3 even if you don't. Their dependency is yours too.

You don't need to eliminate these dependencies (S3 is still very reliable), but you should be able to list them, and each one should fail visibly and gracefully rather than mysteriously.

Down versus degraded

For most affected sites, the homepage returned 200 OK the whole time. A basic uptime check saw nothing wrong while signups failed because avatar uploads errored, checkouts failed because invoice PDFs couldn't be written, and dashboards rendered blank.

Hard-down is the easy case; any monitor catches it. Degradation is where revenue leaks, and catching it means checking features, not just the front door:

  • Keyword checks. Assert that the page contains what a working page contains, so a 200 with an error message counts as down.
  • Transaction-critical endpoints. Monitor the API route behind checkout or upload, not just the marketing page in front of it.
  • The dependency itself. A small probe object fetched from your bucket tells you the storage path works, well before support tickets do.

A reality check on multi-region

The standard post-outage advice, "just go multi-region", deserves some skepticism. Real multi-region operation is expensive, invasive and easy to get subtly wrong, and for most small teams it's the wrong first place to spend money. A more realistic approach:

  • Know your blast radius. Which features die if us-east-1 does? Deciding to accept four hours of broken uploads every few years is legitimate. The point is to decide it in advance.
  • Degrade on purpose. Cache what you can, queue writes for retry, and show "uploads are temporarily unavailable" instead of a stack trace.
  • Put cheap redundancy where it counts. That brings us back to the status page, the one component where independence costs almost nothing.

Know first, and have somewhere to say it

The teams that came out of last Tuesday looking good knew about the problem before their customers told them, and had a working place to acknowledge it. Both are cheap to arrange in advance and painful to improvise.

CompleteStatus handles the first part: external uptime and keyword checks that catch degradation as well as hard-down, with alerts by email, Slack or webhook when a check fails, run from infrastructure independent of yours. For the second, the status pages guide covers setting up a page that doesn't share fate with your stack. Last week we wrote about inheriting your CDN's bugs; this week the same lesson applied to storage. Create a free account and hear about the next us-east-1 afternoon from your monitor, not from Twitter.

Written with AI assistance and reviewed by the CompleteStatus team.