Skip to main content

The GitLab.com Database Incident: A Backup You Haven't Restored Is Just a Hope

Last week GitLab.com lost hours of production data and found its backups weren't working. What happened, and how to check your own before you need them.

The CompleteStatus Team 7 min read

The GitLab.com Database Incident: A Backup You Haven't Restored Is Just a Hope

Want this checked continuously?

CompleteStatus grades your headers, SSL and email security 24/7 — free, commercial use allowed.

Monitor your site free

Last Tuesday evening, GitLab.com went down, and over the following day a lot of engineers watched a recovery happen in public. GitLab kept a running incident document open to everyone, and streamed part of the restore live. It's one of the most transparent outages we've seen from a company of that size, and it deserves credit for that.

It's also uncomfortable reading, because the failure at the heart of it is one that almost every team is exposed to. GitLab has promised a full postmortem and it isn't out yet, so some details below may be revised. What they've published so far is already enough to learn from.

What happened, as far as we know

Based on GitLab's own notes and blog updates:

  • The site was under unusual load, reportedly from spam activity, and database replication from the primary PostgreSQL server to the secondary fell behind and then stopped.
  • While working late to rebuild replication, an engineer ran a command to clear the data directory on what they believed was the secondary. It was the primary. Around 300 GB of production data was deleted before the command was stopped.
  • When the team went to restore, they found that the backup mechanisms they expected to rely on weren't producing usable backups. Regular database dumps had been failing, apparently because of a version mismatch between the dump tool and the database server, so they produced files that were far too small. The failure notifications were being sent by email and, according to GitLab's notes, were being rejected by the receiving side, so nobody saw them.
  • Other options, such as cloud disk snapshots, turned out not to be enabled for the database servers.
  • What saved them was a copy of the database taken about six hours earlier, more or less by chance, to refresh a staging environment. They restored from that, over many hours, partly because the disks holding the copy were slow.

The result, per GitLab: roughly six hours of database changes were lost. That means issues, merge requests, comments and user accounts created in that window. Git repositories and wikis weren't affected, because they're stored separately from the database.

The rm isn't the lesson

It's tempting to focus on the command. Someone typed a destructive command into the wrong terminal late at night after a long day. That will happen to someone on your team, eventually, and there are good mitigations: coloured prompts that differ between production hosts, requiring a second person for destructive operations on primaries, tooling that refuses to run against the node currently serving writes.

But the command is the trigger, not the cause. Operators make mistakes; the system is supposed to survive them. The reason this became a data loss incident rather than a bad hour is that there were several independent safety nets, and when they were needed, most of them turned out not to be there. GitLab's engineers were refreshingly blunt about this in their notes.

Why backups fail silently

The GitLab backup failures are a textbook case of the three ways backups break without anyone noticing.

The job runs but produces garbage. A dump tool that exits early, or writes an error message into the output file, still leaves a file behind. A cron job that "ran" isn't a backup that succeeded. The only honest test of a backup is whether it restores.

The failure alert goes nowhere. Cron's default behaviour is to email output to the job's owner. That email passes through mail servers, spam filters and authentication checks like DMARC, any of which can drop it quietly. An alerting path you haven't tested is exactly as reliable as a backup you haven't tested.

The safety net was never switched on. Features like disk snapshots are often assumed to be on because they're available. Availability and configuration are different things.

What all three have in common: they fail by absence. Nothing breaks. There's just no backup, and no signal that there's no backup. Monitoring built around "alert me when something goes wrong" doesn't help, because nothing visibly goes wrong.

Checks worth doing this week

Here's what we'd suggest, in rough order of value:

  1. Restore something. Take last night's database backup and restore it to a scratch server. Time it. Check the row counts on a few large tables against production. If you've never done this, expect it to take longer than you think, and expect to learn something.
  2. Check backup sizes over time. A backup that suddenly shrinks is a red flag. Even a crude check helps:
#!/bin/sh
# fail loudly if last night's dump is missing or suspiciously small
f=$(ls -t /backups/db-*.sql.gz 2>/dev/null | head -1)
[ -n "$f" ] || { echo "no backup found"; exit 1; }
size=$(stat -c %s "$f")
[ "$size" -gt 500000000 ] || { echo "backup $f is only $size bytes"; exit 1; }
  1. Make backup success report in, rather than failure report out. Instead of relying on an error email, have the backup job actively check in with something when it finishes successfully, and alert when that check-in doesn't arrive on time. This is sometimes called a dead man's switch. It turns silence into a signal, which is the whole problem with backups.
  2. Match tool versions to server versions. If your dump tool and your database server come from different packages, pin them together and test after every upgrade. This is exactly the kind of mismatch that appears after a routine upgrade and never announces itself.
  3. Keep a copy somewhere else. Backups on the same server, the same disk array or the same cloud account share failure modes with production. At least one copy should live somewhere a single mistake or compromise can't reach.
  4. Write the restore runbook, then have someone else follow it. The person who set up backups can usually restore them. The test is whether the on-call engineer at 3 a.m. can.

Replication is not a backup

One more point that GitLab's incident illustrates well. Replication protects you against hardware failure: if the primary's disk dies, the replica has the data. It doesn't protect you against mistakes. A bad DELETE, a broken migration, or a destructive command replicates faithfully to every copy within seconds. You need point-in-time backups for that, kept long enough that you can go back to before the mistake happened.

In GitLab's case, replication had already stopped, which is what started the whole evening. But even if it had been working perfectly, the replica would not have been a defence against the data being deleted on the primary.

Credit where it's due

It would have been easy for GitLab to publish a short apology and move on. Instead, they wrote down every failure in public while it was still embarrassing. That openness is why the whole industry can learn from this one, and it's the same argument we'll keep making: people forgive outages much more readily than they forgive silence.

Outside-in monitoring wouldn't have saved GitLab's data. It tells you when your site is down or slow, which they already knew. The part that applies to all of us is the underlying idea: don't trust a system to report its own failure. Check from somewhere else, and check for the presence of success, not just the absence of errors. That's the approach we take at CompleteStatus, and it's a good rule for backups too.

Written with AI assistance and reviewed by the CompleteStatus team.