Skip to main content

GitHub's Day-Long Outage: Choosing Consistency Over Availability

GitHub spent about a day degraded this weekend and chose data integrity over a fast recovery. What that trade-off means, and what to check in your own pipeline.

The CompleteStatus Team 6 min read

GitHub's Day-Long Outage: Choosing Consistency Over Availability

Want this checked continuously?

CompleteStatus grades your headers, SSL and email security 24/7 — free, commercial use allowed.

Monitor your site free

For roughly a day, from Sunday evening UTC until Monday evening, GitHub was degraded. Git operations mostly worked for much of it, but the website showed stale or inconsistent data, webhooks and GitHub Pages builds were queued rather than delivered, and a lot of teams discovered how much of their daily work runs through one company's servers.

GitHub has published a preliminary update and promised a full analysis. What it says so far: a brief loss of network connectivity between parts of its infrastructure led to problems in its MySQL database topology, and the recovery took as long as it did because GitHub decided to protect data integrity rather than restore service as quickly as possible. We'll wait for the full write-up before drawing conclusions about the specifics. The broad shape is already worth talking about, because it's a trade-off every team with a database eventually faces.

The choice GitHub made

When a replicated database gets split by a network problem, two sides can each accept writes that the other doesn't know about. When the network heals, you have two diverging histories. At that point there are roughly two options:

  1. Get back up fast. Pick one side, discard or paper over the conflicting writes, and resume. Service returns sooner. Some data may be lost or inconsistent, and you may not know exactly which.
  2. Reconcile first. Keep the site degraded or read-only while you restore from backups, replay the writes that happened on each side and resolve conflicts. Service returns much later. No user data is silently thrown away.

GitHub's update says it chose the second. For a service that holds code, issues and review history for millions of projects, that seems defensible: an hour of downtime annoys people, while a silently lost merge or comment can damage their trust in the data forever. But it meant a day of degraded service rather than an hour of outage, and it's worth being honest that both options have real costs.

This is the practical version of the trade-off the CAP theorem describes: during a network partition, a distributed system has to choose between staying available and staying consistent. Most of the time you don't feel it. When you do, you find out whether anyone decided in advance which one your system prefers.

Decide before the partition, not during it

The time to answer "availability or consistency?" is not 11 p.m. on a Sunday with the site half down. It's worth writing down for each important data store:

  • What happens if the primary is cut off from its replicas for 30 seconds? Does something promote a replica automatically? Can the old primary keep accepting writes while that happens?
  • If both sides accept writes, how would you know? Is there a way to detect divergence, or would you find out from customer reports?
  • Which data can you afford to lose? Page view counters and session data may be fine to drop. Orders, payments and user content usually aren't. The answer can be different per table.
  • How long does a full restore take? Not in theory. Measured. If you haven't restored recently, you don't know. (GitLab's database incident last year is the standard reference for what happens when you find out during the emergency.)

Automatic failover is designed to make outages shorter, and usually it does. But the partition scenarios where it guesses wrong are exactly the ones that make recovery long and complicated. Knowing how your failover behaves on a short, messy network blip is more useful than knowing how it behaves in a clean server crash.

"Degraded" is hard to monitor and hard to explain

A lot of this outage wasn't "down" in the simple sense. Pages loaded. Some data was out of date. Some features worked and others didn't. That's a difficult state for monitoring and for communication.

For monitoring, a check that only asks "does the home page return 200?" would have been green for much of the day. The problems were in freshness and in background processing: webhooks not firing, builds not running. If your own service can fail in that way, think about checks that cover it:

  • A check that reads something your system writes regularly and alerts if it's too old.
  • A check that your background queues are actually draining, not just that the workers are running.
  • A synthetic end-to-end action, such as creating a test record and reading it back, rather than just loading a page.

For communication, GitHub posted regular updates on its status page throughout, explaining what was working, what wasn't and why recovery was slow. That matters. A status page that sits at "investigating" for hours invites people to assume the worst. As we argued after the S3 outage last year, the status page is the one piece of your infrastructure that has to keep working when everything else doesn't, and it has to say something useful.

Your pipeline depends on GitHub more than you think

The second lesson is about the people downstream. Plenty of teams found on Monday that they couldn't deploy, because their deploy process pulled code from GitHub, or waited for a webhook to trigger CI, or installed dependencies from a GitHub URL.

Some questions to ask about your own setup:

  • Can you deploy if GitHub is unavailable? If there's a production fix to ship during an outage like this, is there a documented way to build from a local clone and deploy directly?
  • What happens to webhooks you missed? If your CI or chat integrations rely on webhooks, do they catch up when delivery resumes, or do events just vanish? Is there a way to trigger a build manually?
  • Are dependencies pinned and cached? Package installs that fetch from GitHub at build time will fail when it's degraded. A local cache or mirror removes that dependency for builds.
  • Do you monitor your dependencies' status? Knowing within minutes that GitHub, your payment provider or your DNS host is having problems saves a lot of time spent debugging your own systems. We've made this point before about DDoS attacks against GitHub and about DNS providers: your availability includes theirs.

Wait for the postmortem

GitHub's full analysis should add a lot of detail, and some of what's reported today may turn out to be incomplete. It will be worth reading when it arrives, especially the part about why automated failover behaved as it did and what they plan to change.

For now, two things to take away. First, decide in advance how your databases should behave during a partition, and test a restore so you know how long "reconcile first" would take. Second, check how much of your deploy process assumes a third party is up. External monitoring of your own endpoints, plus a watch on the services you depend on, is how you find out quickly which side of that line a problem is on. That's the kind of monitoring CompleteStatus is built to do.

Written with AI assistance and reviewed by the CompleteStatus team.