Skip to main content

Yesterday's BGP Route Leak: How a Small ISP Rerouted the Internet

A Pennsylvania ISP's route optimizer leaked routes, Verizon propagated them, and traffic for Cloudflare and Amazon went sideways. What BGP teaches us.

The CompleteStatus Team 6 min read

Yesterday's BGP Route Leak: How a Small ISP Rerouted the Internet

Want this checked continuously?

CompleteStatus grades your headers, SSL and email security 24/7 — free, commercial use allowed.

Monitor your site free

Yesterday morning, a sizeable part of the internet was routed through a metals company in Pennsylvania. That's not a metaphor. For roughly two hours starting around 10:30 UTC on June 24, traffic destined for Cloudflare, Amazon, Facebook and others was sent through networks that had no business carrying it, and in many cases couldn't, so the traffic was dropped. Cloudflare says it lost around 15% of its global traffic at the peak. Sites behind the affected networks were unreachable for many users and fine for others.

Nobody attacked anything. The internet's routing system did what it was built to do, and that's the problem. If you run anything on the web, this is worth understanding, because it's a kind of outage that starts entirely outside your infrastructure and your control.

What happened

The sequence, as pieced together from routing data and the reports published so far:

  • A small ISP ran a BGP route optimizer. DQE Communications, a regional provider in Pennsylvania, used software that splits routes into smaller, more specific prefixes to steer traffic over preferred internal paths. Those synthetic routes are meant to stay inside the network.
  • They leaked to a customer. The optimized routes were announced to one of DQE's customers, Allegheny Technologies, a specialty metals company, whose network passed them on to its other upstream provider.
  • That upstream was Verizon, and Verizon accepted them. This turned a local mistake into a global incident. Verizon, one of the largest transit networks in the world, propagated the leaked routes to the wider internet without filtering them.
  • More specific wins. Given two routes to the same destination, BGP routers prefer the more specific prefix. The leaked routes were more specific than the legitimate ones, so routers across the internet sent traffic for Cloudflare, Amazon and others down a path that ran through a metals company's network connection. That link was overwhelmed almost immediately.

About two hours later the leak was withdrawn and routes converged back to normal. For affected users, it was a complete outage for that time.

BGP still runs on trust

The Border Gateway Protocol is how the internet's roughly 65,000 independent networks tell each other which addresses they can reach. It comes from an era when the internet was small enough that everyone more or less knew everyone, and at the protocol level announcements are believed. If a network says it has a route to some addresses, its neighbors accept that unless they've configured filters saying otherwise.

Those filters are supposed to be the safety net: each network vetting what it accepts from customers and peers, using registries of who legitimately holds which prefixes. Small networks get this wrong regularly, and the damage stays small. Yesterday was different because a tier-1 transit provider accepted and re-announced thousands of leaked, more-specific routes from a customer. Filtering at that boundary (even a simple prefix limit would have helped) is basic practice for a network of that size, and it evidently wasn't in place.

There is a cryptographic fix in progress. RPKI (Resource Public Key Infrastructure) lets address holders sign statements about which networks may originate their prefixes, and lets routers reject invalid announcements automatically. Cloudflare and a few others have been vocal adopters, and some large carriers, including AT&T, have started dropping RPKI-invalid routes this year. But adoption is still thin: a minority of routes are signed, and fewer networks validate. Until that changes, incidents like yesterday's are a structural feature of the internet rather than a freak event.

The monitoring lesson: "down" is not one thing

This matters even if you never touch a router. During yesterday's leak, a site behind an affected network was, at the same time:

  • Completely unreachable for users whose traffic followed the bad routes,
  • Fine for users whose providers never accepted the leak, and
  • Fine from inside its own datacenter: servers up, dashboards green, load balancers passing health checks.

Every internal signal said healthy, because nothing internal was broken. The failure was in the path between users and infrastructure, and the only way to see a path failure is to check from where your users are, in several places at once.

That's the strongest argument for multi-location external monitoring:

  • One probe location can't tell you much. If your single monitoring node is on a network that accepted the leaked routes, you get a false "everything is down". If it's on a clean network, you get a false "everything is fine".
  • Multiple locations turn confusion into diagnosis. Three of five regions failing to reach you while two succeed is a distinctive pattern that points to routing, not your server. That changes your response: you're checking BGP looking glasses and your providers' status pages instead of restarting services that were never broken.
  • It changes what you tell customers. "Some networks can't currently reach us because of an internet routing incident; our systems are operational" is a very different message from a generic apology, and you can only send it if your monitoring can tell the two cases apart.
# The view from one network is not the view from the internet.
# Yesterday, these could disagree for two hours:
curl -sI https://yoursite.com   # from your office: 200 OK
curl -sI https://yoursite.com   # from a leaked-route network: timeout

What you can do

You can't fix BGP from your web app, but you have some options:

  • Monitor from multiple, geographically distinct locations, and pay attention to partial failure patterns, not just up and down.
  • Ask your providers about routing hygiene. Do your transit and hosting providers filter customer announcements? Have they deployed RPKI origin validation? These are reasonable questions to ask vendors.
  • Have a status page and a communication plan for outages that aren't your fault. Your users can't tell the difference, and "not our fault" without evidence reads as an excuse.
  • Keep the timeline. Multi-location check history gives you the start, end and geographic scope of an incident, which is what you need for the postmortem and any SLA conversation.

Checking from where your users are

Yesterday's leak was fixed by engineers at other companies, on infrastructure you'll never touch, and the next one will be too. What you control is knowing quickly that it's happening, where it's visible from, and whether the problem is your stack or the network between you and your users. CompleteStatus runs your uptime checks from multiple locations and shows which of them can and can't reach you, with alerts when a failure is confirmed. See what it checks, or create a free account.

Written with AI assistance and reviewed by the CompleteStatus team.