The Rogers Outage: When One Carrier Takes Down Phones, Payments and 911
Friday's Rogers outage knocked out mobile, internet and Interac debit across Canada. What it says about single-carrier risk and where your alerts go.
Want this checked continuously?
CompleteStatus grades your headers, SSL and email security 24/7 — free, commercial use allowed.
Early last Friday morning, July 8, Rogers Communications, one of Canada's largest telecom carriers, dropped off the internet. Not a region, not a service: wireless, home internet, business connectivity and the Fido and Chatr brands that ride on the same network all went down together, starting around 4:45 a.m. Eastern. For many customers service didn't come back until late Friday or into Saturday, and some reported problems lingering over the weekend.
The knock-on effects are what turned a carrier outage into a national story. Interac's debit network, which Canadian shoppers use for a huge share of in-store card payments, went down across the country because it depended on Rogers. Shops put up "cash only" signs. Some people couldn't reach 911 from Rogers phones. Bank call centres, government offices and hospitals reported problems with phone lines. Plenty of businesses that had never thought of themselves as "Rogers customers" discovered that something they depended on was.
We don't have a full postmortem yet, and we won't pretend to. What Rogers has said publicly is that a maintenance update in its core network caused some of its routers to malfunction. Outside observers, including Cloudflare's network data, saw large parts of Rogers' address space effectively vanish from global routing while it was happening, which fits the "the whole network went away at once" experience customers had. The detail will come later. The lessons for everyone else are already visible.
The dependency you didn't write down
If you asked most Canadian retailers on Thursday what their payment system depended on, they would have said their card terminal vendor and their bank. Very few would have said "one specific telecom carrier's core network." That's the pattern in almost every large outage we've written about, from the Dyn DNS attack to the Fastly outage last June: the thing that fails is two or three layers below the thing you think you depend on, and it's shared by far more people than you'd guess.
A useful exercise this week is to trace your critical paths all the way down. For each thing that has to work for your business to operate, ask:
- What network does it actually leave the building on? One ISP? One mobile carrier for the backup LTE modem?
- Does the backup link share anything with the primary? Same carrier under a different brand, same fibre conduit, same upstream?
- Which of your vendors has a single carrier in their path? You can't always find out, but you can ask, especially for payments and phones.
"We have two internet connections" is worth much less if both are resold from the same backbone. After Friday, a lot of backup plans in Canada are going to be re-examined for exactly that reason.
Your monitoring can't alert over the thing that's down
An uncomfortable operational question: if Rogers had been your carrier, how would you have found out your systems were in trouble at 5 a.m.?
A lot of alerting paths quietly assume the network is fine:
- SMS and phone-call alerts go to phones. If the on-call engineer's phone is on the affected carrier, the page never arrives. Nobody gets an error; the message just doesn't land.
- Internal monitoring running inside the affected network can see that things are wrong, but it can't tell anyone, because its only route out is the link that failed.
- Chat-based alerts need the recipient to have connectivity. Home internet and mobile on the same carrier (bundles are popular for a reason) means both are gone at once.
None of this is exotic. It's the same fate-sharing problem as hosting your status page on your own infrastructure, which we wrote about after the S3 outage, applied to people instead of servers.
Practical fixes, most of them cheap:
- Monitor from outside. Checks that run from networks you don't operate can still see, and still report, when your own network is unreachable. Internal monitoring is valuable for depth; it's useless for "is anything reaching us at all."
- Send critical alerts through at least two unrelated channels. Email plus a chat integration plus SMS is not three channels if all three terminate on one phone on one carrier. Make sure at least one person on the escalation path is on a different carrier, or reachable on a different network.
- Write down the out-of-band plan. Who calls whom, on what, when the normal tools are dead. A laminated card with phone numbers sounds quaint until the day it's the only thing that works.
- Test it. Pull the plug on the office uplink during a quiet hour and see which alerts actually arrive, and where.
Degraded modes beat heroics
The businesses that coped best on Friday weren't the ones with the cleverest networks; they were the ones with a plan for operating badly. Shops that could take cash or cheques, or had a terminal on a different carrier, kept trading. Clinics that could fall back to paper kept seeing patients.
For web and SaaS teams, the equivalent is deciding in advance what "degraded" looks like:
- If your payment provider is unreachable, can you queue orders and charge later, or show a clear message instead of a spinning checkout?
- If a third-party API is down, does your page fail gracefully or throw a 500 at the whole request?
- If a region or provider disappears, is failover automatic, or does it need someone to log in to a console that may itself be unreachable?
These are design decisions, and they're far easier to make on a Tuesday afternoon than during an outage.
Communication during a total outage
One more lesson from Friday: Rogers' customers had a hard time finding out what was going on because their connectivity was the thing that was broken. People drove to find Wi-Fi from another provider just to read an update. Updates on social media only reach people who can get online.
If you run a service, think about where your customers will look for news when your service is down, and make sure your updates are there: a status page hosted independently of your own infrastructure, an email list, a social account. And be specific. "We are aware of an issue affecting some customers" tells nobody anything; "all mobile and internet service is down nationally, 911 calls may not connect, use another phone if you can" is information people can act on.
Regulators are paying attention
Canada's industry minister met the heads of the major carriers on Monday and has asked them to put agreements in place, reportedly within 60 days, covering emergency roaming, mutual assistance during outages and better communication with the public. The CRTC is also expected to demand a detailed account from Rogers. Expect the conversation about carrier diversity for critical services, including payments and emergency calling, to run for months.
Most of us don't run national infrastructure, but the same question applies at every scale: when the network under you disappears, what still works, and who finds out?
Check your own blast radius
Take thirty minutes this week and write down the network path for your three most important services and your alerting. If any of them collapse to a single carrier, you have found your Friday. External checks, whether you run them yourself or use a service like CompleteStatus, are the part of the answer that keeps working when your own network doesn't, provided the alerts go somewhere that's still reachable.
Written with AI assistance and reviewed by the CompleteStatus team.
Get notified when CompleteStatus opens
New accounts are closed while we're in private beta. Leave your email and we'll send one message the moment sign-ups open — nothing else.