Sunday's Google Cloud Outage: When Your Provider's Network Has a Bad Afternoon
A Google configuration change on June 2 congested networking across US regions for hours. What it teaches about shared fate, alert paths and outside-in checks.
Want this checked continuously?
CompleteStatus grades your headers, SSL and email security 24/7 — free, commercial use allowed.
On Sunday afternoon, US time, a large slice of Google got very slow and then, for many users, stopped working. YouTube buffered, Gmail and G Suite timed out, and Google Cloud customers, including consumer apps like Snapchat and Discord according to widespread user reports, had trouble serving anyone in the eastern United States. The disruption started a little before noon Pacific and lasted around four hours before things were back to normal.
Google hasn't published a full postmortem yet, but it did post an early explanation on Monday. The short version: a configuration change intended for a small number of servers in a single region was applied to a larger number of servers across several neighbouring regions. Those regions then stopped using more than half of their available network capacity. The network didn't go down. It went into severe congestion, and Google's traffic management then had to decide what got through.
Google's account also includes a detail every on-call engineer will recognise: the congestion slowed down the tools its own engineers needed to diagnose and fix the problem. We'll come back to that.
What "the cloud was down" actually meant
The failure didn't look the same for everyone, which is what made it confusing to watch:
- Some requests worked, slowly. Congestion drops a share of packets rather than all of them. Pages loaded in 20 seconds instead of 1, or loaded without images, or failed on the third API call of five.
- Some regions were fine. Workloads outside the affected area kept running. Plenty of Google Cloud customers saw nothing at all.
- Internal health looked healthy. Instances were up, CPUs were idle, processes were running. A VM checking its own web server on localhost got a fast 200 the whole time.
- Google's services were hit too. Early figures from Google say YouTube saw a drop of around 10% in global views, and Cloud Storage traffic fell by roughly 30%. When the provider's own flagship products are struggling, you know it isn't just your app.
If your monitoring only answers "is the process alive?", Sunday was a green afternoon. If it answers "can a user somewhere else load this page in a reasonable time?", Sunday was a four-hour incident. Only the second one matches what your customers experienced.
Regions share more than you think
Cloud providers sell regions as independent failure domains, and for most failures (power, hardware, a bad rack, a lost data hall) they are. Sunday is a reminder that regions are still run by one company, with one set of configuration systems, one network backbone and one set of automation that can push a change to the wrong place.
That doesn't make multi-region architecture pointless. It means the failures it protects you from are specific, and a provider-wide control problem isn't one of them. When you write down your availability assumptions, be honest about which layer they rely on:
- Multi-zone protects you against losing a building or a rack.
- Multi-region protects you against losing a region's infrastructure.
- Neither protects you against the provider's shared network or management plane misbehaving. Only a second provider (or a static fallback hosted elsewhere) does that.
Most teams will reasonably decide a second provider isn't worth the cost and complexity. That's a fine decision, as long as it's a decision and not an assumption.
The tools you fix things with must survive the thing that broke
Google's engineers were slowed down by the same congestion they were trying to fix. Scale that down to a small team and it's a very familiar failure:
- The alerting system runs on the same cloud project as the app, so it's degraded at exactly the wrong moment.
- The alert emails go to a Google-hosted inbox. On Sunday, some of those inboxes were among the things timing out.
- The runbook lives in a wiki on the same infrastructure.
- The status page is served from the same load balancer as the product. (We wrote about this after the S3 outage in 2017, and it keeps being true.)
- The only way to reach the servers is through a bastion host in the affected region.
None of these are exotic. They accumulate because putting everything in one place is the default, and the default is convenient right up until it isn't. A short audit is worth doing this week:
- Where does your monitoring run, and would it notice if your provider's network degraded?
- How do alerts reach a human? Is there at least one path (SMS, a phone call, a chat service on a different provider) that doesn't depend on the same company as your production?
- Can you read your runbook and update your status page from a phone on mobile data, with your primary cloud completely unreachable?
- Do you have a way into your infrastructure that isn't in the same region as the failure?
Watch for slow, not just down
Congestion incidents are a strong argument for alerting on response time as well as availability. A check that only fails on a timeout of 30 seconds will flap on Sunday, succeeding sometimes and failing sometimes, and if it needs two consecutive failures to alert, it may never alert at all.
A better setup for this kind of failure:
- Record response time on every check and alert on sustained degradation (for example, median over 3 seconds for 10 minutes), not only on hard failures.
- Alert on error rate across a window, not only on consecutive failures. Three failures out of the last six checks is an incident even if they were never back to back.
- Check from more than one network. A check running in the same cloud region as your app shares its fate. One running from a different provider and a different geography sees what your users see.
A quick way to see the difference yourself is to time the same request from two places and compare:
curl -s -o /dev/null -w "dns=%{time_namelookup} connect=%{time_connect} tls=%{time_appconnect} total=%{time_total}\n" https://yoursite.example/
During a congestion event, connect and total blow out while the server-side logs show nothing unusual, because the time is being lost before requests ever arrive.
Talk about it the right way
When your provider is the cause, it's tempting to say nothing and wait for it to pass. Customers don't care whose fault it is, but they do care whether you know. "Our hosting provider is experiencing a network incident affecting the eastern US; our systems are otherwise healthy and we're monitoring" is a perfectly good update, and far better than silence. It's only possible if your monitoring can tell you it's the network and not your code, and if your status page is somewhere the incident can't reach.
The takeaway
Sunday's outage wasn't caused by anything Google's customers did, and there wasn't much they could do during it except wait. But the difference between teams that had a calm afternoon and teams that had a bad one was mostly decided beforehand: whether monitoring watched from outside, whether alerts had a path that didn't share fate with production, and whether anyone could say something accurate to customers within the first fifteen minutes.
That's the part worth fixing before the next one. External checks, from a network that isn't your provider's, are what CompleteStatus is built around, and they're a sensible first step whichever tool you use.
Written with AI assistance and reviewed by the CompleteStatus team.
Get notified when CompleteStatus opens
New accounts are closed while we're in private beta. Leave your email and we'll send one message the moment sign-ups open — nothing else.