The OVHcloud Strasbourg Fire: Were Your Backups in the Same Building?
Last week's fire destroyed a data centre in Strasbourg and took backups with it. How to check where your copies really live, and how fast you could restore.
Want this checked continuously?
CompleteStatus grades your headers, SSL and email security 24/7 — free, commercial use allowed.
In the early hours of Wednesday, March 10, a fire broke out at OVHcloud's data centre campus in Strasbourg. By morning, one building, SBG2, had been destroyed. Part of the neighbouring SBG1 was damaged, and the other two buildings on the site, SBG3 and SBG4, were shut down as a precaution and have stayed offline while OVHcloud works through restarting them. Nobody was hurt, which is the most important fact about it.
For customers, the consequences are still unfolding a week later. Netcraft estimated that millions of websites were served from IP addresses in the affected data centres. Some customers will be back online as SBG1, SBG3 and SBG4 come back. Some lost servers entirely. And a number of customers have said publicly that they lost not only their servers but their backups, because the backups were stored on the same site.
The cause of the fire hasn't been established yet. OVHcloud's founder has been giving regular updates, and there's been early discussion of the building's power equipment, but an investigation is under way and it's too early to say anything definitive. This post isn't about the cause. It's about the question every team should ask this week, whether they use OVHcloud or not: if our building burned down tonight, what would we have left?
"A backup" is not a location
Most teams can answer "do you have backups?" with a confident yes. Far fewer can answer "where, physically, is the most recent copy?" without looking it up. Common answers, once people check:
- On the same server, in a
/backupsdirectory. This protects against someone deleting a table. It protects against nothing that happens to the machine. - On another server in the same data centre. Better against a disk failure. No help at all against fire, flood, or a power incident that takes out the building.
- In the provider's backup product, which, depending on the product and the settings, may be in the same facility as the server. Worth reading the documentation carefully rather than assuming.
- In object storage "in the same region". Regions and availability zones are real engineering, but the definition varies by provider. Some regions are several buildings kilometres apart; some are several buildings on one campus.
Strasbourg was one campus with four buildings. For a customer whose server was in SBG2 and whose backup was in SBG1, a "different data centre" wasn't far enough.
The 3-2-1 rule, taken literally
The old rule still holds up: three copies of your data, on two different types of storage, with one copy off-site. This week is a good reminder to be strict about the last part:
- Off-site means a different facility, far enough away that a single physical event can't reach both. A different city is a sensible minimum. A different provider is better still, because it also protects you against account problems, billing disputes and provider-wide outages.
- Off-site means your credentials can't delete it. If an attacker, or a buggy script, can delete production and the backups with the same key, you don't have two copies. Use separate credentials, write-only access from production, or object versioning and retention locks where your storage supports them.
- Off-site copies need monitoring too. A nightly job that has silently failed since November is not a backup. Alert when the newest file in the off-site bucket is older than it should be, and when it's suspiciously small.
A crude but useful check you can schedule today:
# Alert if the newest off-site backup is older than 26 hours
latest=$(aws s3 ls s3://example-offsite-backups/db/ | sort | tail -n 1 | awk '{print $1" "$2}')
age=$(( $(date +%s) - $(date -d "$latest" +%s) ))
[ "$age" -gt 93600 ] && echo "Off-site backup is stale: $latest" >&2 && exit 1
Swap in your own storage tool and alerting. The point is that something other than a human checks it every day.
Restore time is the number that matters
Having the data is step one. The question customers actually care about is how long until they're back. Teams that had good off-site backups last week still faced a long day or several, because restoring means:
- New servers. Provisioning replacements, possibly with a different provider or in a different location, and possibly competing with every other affected customer for the same capacity.
- Configuration. Web server config, TLS certificates, cron jobs, firewall rules, environment variables. If these only existed on the server that burned, they have to be recreated from memory.
- Data. Downloading and importing backups, which for a large database can take hours by itself.
- DNS. Pointing traffic at the new location. If your TTLs are a day long, some visitors will be trying the dead IP address until tomorrow.
Each step is a place where an untested plan turns into improvisation. The only honest way to know your restore time is to do a restore: pick a quiet afternoon, build the site from backups on a fresh server somewhere else, and time it. The first attempt usually turns up a missing config file, a password nobody recorded, or a backup that doesn't include what everyone assumed. Better to find that on a Tuesday afternoon than during an incident.
Don't keep your recovery tools in the same place
The same logic applies to everything you'd need during recovery:
- DNS. If your DNS is hosted by the same provider as your servers, a provider-level problem can stop you repointing traffic. We covered choosing a DNS provider with redundancy in mind in 2019, and it applies here.
- Status page. Customers need to hear from you while the site is down, so the page that tells them can't be on the servers that are down. The S3 outage in 2017 taught that lesson in a different way.
- Documentation and secrets. The runbook, the list of servers, and the credentials to rebuild them need to be reachable when your infrastructure isn't.
- Monitoring. Something outside the building needs to notice the building is gone and tell you. At 1am, that's what gets you out of bed.
A checklist for this week
- Write down, for each production system, the physical location of the server and of every backup copy.
- Make sure at least one copy is in a different city, ideally with a different provider, under separate credentials.
- Add alerting for stale or undersized backups.
- Lower DNS TTLs on records you might need to change in an emergency (five minutes to an hour is reasonable for most).
- Schedule a real restore test and time it.
- Check that your status page, DNS, runbooks and monitoring would all survive losing your primary provider.
Fires in data centres are rare, and the industry puts a great deal of effort into preventing them. But "rare" applied across thousands of facilities and many years means it will happen to someone, and last week it happened to a lot of people at once. CompleteStatus checks from outside your infrastructure, which is where the monitor needs to be when the infrastructure itself is the thing that's gone. Whatever you use, make sure it isn't in the same building as everything else.
Written with AI assistance and reviewed by the CompleteStatus team.
Get notified when CompleteStatus opens
New accounts are closed while we're in private beta. Leave your email and we'll send one message the moment sign-ups open — nothing else.