One Bad Regex: Lessons From the July 2 Cloudflare Outage
A WAF rule with a backtracking regex pushed Cloudflare's edge CPUs to 100% and served 502s worldwide. What the outage teaches every team.
Want this checked continuously?
CompleteStatus grades your headers, SSL and email security 24/7 — free, commercial use allowed.
Last Tuesday, July 2, a very large part of the web disappeared behind 502 errors, everywhere at once. For roughly half an hour, sites behind Cloudflare returned "502 Bad Gateway" to visitors while their own origin servers sat idle and healthy. The cause wasn't an attack, a fiber cut or a datacenter fire. It was a single new rule in Cloudflare's Web Application Firewall containing a regular expression that backtracked badly enough to use 100% of the CPU on machines across Cloudflare's global edge.
Two weeks ago it was a BGP route leak sending traffic the wrong way; this time the cause was a config deploy. Cloudflare published an explanation the same day that named the cause plainly, and has said a more detailed technical write-up will follow. Even with only the initial account, the failure mode is clear enough to learn from.
What happened, briefly
Cloudflare regularly ships new WAF rules to respond to emerging threats, and the pipeline for those rules is deliberately fast, because blocking a new exploit is time-sensitive. On July 2, a new rule contained a regular expression that backtracked excessively. It was deployed globally in one step, CPU on the edge servers went to 100%, and the proxy layer that serves customer traffic was starved. 502s appeared worldwide within minutes. Cloudflare disabled the WAF managed rules globally, and traffic recovered in well under an hour.
Each part of that deserves its own lesson.
Lesson 1: regex is production code
Cloudflare hasn't yet published the full expression, but the general danger is well known: patterns with several overlapping, greedy quantifiers. A simplified example of the shape:
// A dangerous shape (simplified):
const rule = /.*.*=.*/;
"x=x".match(rule); // fine: matches quickly
("x".repeat(30000)).match(rule); // no "=" anywhere: the engine tries every
// way to split the input between the two
// wildcards before giving up
A backtracking engine tries one way to divide the input among the quantifiers, fails, backs up and tries another. With overlapping wildcards, the number of attempts grows much faster than the input: polynomially for a pattern like this one, exponentially for nested quantifiers like (a+)+. The regex doesn't crash; it just keeps computing at full CPU. On one machine that's a slow request. In a WAF rule evaluated against traffic on every edge server, it's a global outage.
Takeaways for any codebase:
- Review regexes like code, with worst-case inputs in mind, not just the matches you expect.
- Where regexes meet untrusted input, prefer engines with execution guarantees. RE2-style non-backtracking engines run in linear time, and timeouts limit the damage where they can't be used.
- Budget CPU for pattern matching. Anything that evaluates patterns against user-controlled input is a denial-of-service surface. The input doesn't have to be malicious, just unlucky.
Lesson 2: stage every rollout, including config
The regex was the trigger; the deployment model is what made it global. Cloudflare's code changes normally roll out gradually. WAF rules, by Cloudflare's own account, went out on a faster path built for urgent threat response, and that path reached everywhere at once.
At scale, there's no meaningful difference between deploying code and deploying configuration. A WAF rule, a feature flag, a routing weight, a JSON file: if it changes what production machines do, it can take production down, and it deserves the same protections:
- Canary first. One machine, one datacenter or a small percentage of traffic, so a bad change shows itself somewhere you can afford it. Even a few minutes at 1% would likely have made July 2 a minor event.
- Watch a health signal during rollout, automatically. A staged rollout only helps if something is checking the canaries. CPU saturation across a canary group seconds after a rule push is about as loud as a signal gets.
- Don't let the emergency lane become the default. Fast-path deploys exist for good reasons, but every fast path is a bet that the change is safe, and rules written under time pressure are the ones most likely to be wrong.
Lesson 3: kill switches you've actually rehearsed
What ended the outage was a global switch disabling the WAF rules. Having that option is what kept this to around half an hour. It's worth asking whether you have the equivalent:
- Build the big red button before you need it: a fast way to disable any recently shipped subsystem globally.
- Make sure it works while production is on fire. Emergency tooling that depends on the broken system won't help. Keep a management path that doesn't go through your own edge, proxy or auth stack.
- Rehearse. An untested kill switch is a guess, and the middle of a global outage is a bad time to find out it doesn't work.
Lesson 4: recognize a shared-dependency outage
From the outside, July 2 had a clear pattern. If you ran ten sites behind Cloudflare, they didn't fail randomly; they all returned 502s within the same minute, from every check location. That pattern tells you a lot:
- Everything down together, everywhere: a shared dependency (CDN, DNS provider, host). Check their status page, not your servers.
- One site down, others fine: your app or origin. Start debugging.
- Down from some locations, up from others: network path or routing, which calls for a different playbook.
Teams with independent, multi-location monitoring knew within a minute or two that the problem was upstream, told their customers, and waited. Teams without it spent half an hour restarting healthy origin servers. It's also a reminder that monitoring through your CDN monitors the CDN too, which is what you want, provided you can read the pattern when it happens.
Credit where due
Cloudflare's same-day explanation named the root cause and took responsibility, which is the standard every vendor should be held to. Too many incident reports still say "a small subset of customers may have experienced elevated error rates" and leave it there.
Your part is knowing, quickly and independently, when a shared dependency takes you down. CompleteStatus checks your sites from multiple locations and alerts you when confirmed failures appear, so when every monitor fires at once you can recognize the pattern and go straight to the right response. See what it monitors, or create a free account.
Written with AI assistance and reviewed by the CompleteStatus team.
Get notified when CompleteStatus opens
New accounts are closed while we're in private beta. Leave your email and we'll send one message the moment sign-ups open — nothing else.