A regex took down Cloudflare. 27 minutes.
July 2, 2019, 13:42 UTC: one new rule for Cloudflare's Web Application Firewall goes live in 180+ cities in about two seconds.
July 2, 2019, 13:42 UTC: one new rule for Cloudflare's Web Application Firewall goes live in 180+ cities in about two seconds. It is an XSS rule in simulate mode, so it blocks nothing, but it still runs on every request, and it ends in .*(?:.*=.*). PCRE backtracks, every CPU core serving HTTP hits 100 %, and every Cloudflare-proxied site returns 502 for 27 minutes; traffic drops 82 %. The kill switch sits behind Cloudflare Access, which sits behind Cloudflare. Postmortem, from Cloudflare's own write-up: the CPU guard removed weeks earlier in a refactor meant to save CPU, the procedure that let any rule skip staging, the engine with no complexity guarantee, the 11 causes and the 7 fixes. Verdict on the fix: SHIP IT.
Read the written edition (English) ↗
What this video covers
- 13:31–13:42 UTC: PR merged, CI green (no CPU test), Quicksilver pushes the rule to 180+ cities; WAF rules skip the DOG → PIG → Canary stages
- 13:45–14:07: first page; CPU 100 % worldwide, traffic −82 %, 502s everywhere; the internal control panel is behind Cloudflare Access, which is down; some credentials have expired; a rarely drilled bypass
- 14:07–14:09: global WAF terminate; traffic and CPU normal after 27 minutes; 14:52 WAF back minus the rule
- Jul 2 (15:50 UTC) and Jul 12: Cloudflare publishes the same-day note and the full postmortem; CPU guard re-added, 3,868 rules re-read, staged rollouts, move to a linear-time regex engine
Transcript
0:00 One regular expression goes live on every Cloudflare server at once, and for the next twenty-seven minutes the sites behind it are a 502 page, which, for a firewall, is the strictest possible setting. July 2nd, 2019, 13:42 UTC. Cloudflare posts within two hours: not an attack, a bad deploy, traffic down 82 percent. Ten days later CTO John Graham-Cumming publishes the full postmortem, regex included, and Hacker News gives it 698 points,
0:26 which for an outage is a standing ovation. How it happens, why it is possible, and who actually gets the blame. This is The Daily Diff, postmortem. 13:31. A pull request merges: one new firewall rule against cross-site scripting, in simulate mode, so it blocks nothing. 13:37, tests pass; none measures CPU. 13:42, the rule ships to 180 cities in two seconds, because WAF rules skip the dog, pig and canary stages other releases get.
0:54 13:45, the first page. 13:49, Hacker News has a thread on the status page, which still says all systems operational. The rule ends in dot-star, dot-star, equals, dot-star: anything, then anything, then an equals sign. PCRE guesses greedily, fails, and backtracks through every other split. x equals x takes 23 steps. Twenty x's after the equals: 555.
1:15 Twenty x's, no equals sign: 4,067 steps to find nothing. Run that on every request and every core is at one hundred percent, doing nothing, thoroughly. Two guards should have caught it. The CPU limit on rules was removed by mistake weeks earlier, in a refactor meant to make the WAF use less CPU. And the procedure lets any rule skip staging, because rules exist to stop live attacks; this one was not an emergency, and it went global anyway.
1:37 14:00, the WAF is identified; no attack. 14:02, someone proposes the global terminate: one component, off, worldwide. The switch is behind Cloudflare Access. Cloudflare Access is behind Cloudflare. Some credentials have expired from disuse, so the fastest network on the internet spends five minutes on a bypass nobody drilled. 14:07, kill. 14:09, traffic normal. git blame: a rollout with one speed, global; a CPU guard refactored away by
2:04 accident; a regex engine with no upper bound. Not the engineer who wrote the rule: the postmortem lists eleven causes and names nobody. Blast radius: 27 minutes, 82 percent of traffic, 100 percent CPU on every core, in every city. The dashboard and API sit behind the same edge, so customers cannot even switch it off. Hacker News, under the postmortem: they had one problem, used a regular expression, now they have two. Old joke. Still compiles.
2:28 Verdict, postmortem: ship it. The CPU guard is back, all 3,868 rules get read by hand, rules go through staging, and the engine moves to one with linear-time guarantees, published by Ken Thompson in 1968. Monday: no dot-star dot-star in anything that runs per request, and keep the kill switch off the thing it kills. Send me the incident you are still not allowed to talk about, in the comments, or at the daily diff dot dev.
2:53 And that's the diff for today. I'm Niko from Axrisi. Merge responsibly.
Sources
- John Graham-Cumming, "Details of the Cloudflare outage on July 2, 2019" (Jul 12, 2019)blog.cloudflare.com
- Matthew Prince, "Cloudflare outage caused by bad software deploy (updated)" (Jul 2, 2019)blog.cloudflare.com
- Matthew Prince on X, Jul 2, 2019, 14:22 UTCx.com
- Matthew Prince on X, Jul 2, 2019, 14:36 UTC ("No evidence yet attack related")x.com
- Hacker News, Jul 2, 2019, 13:49 UTC — "Cloudflare Network Performance Issues" (631 points)news.ycombinator.com
- Hacker News, Jul 2, 2019 — "Cloudflare outage caused by bad software deploy" (348 points)news.ycombinator.com
- Hacker News, Jul 12, 2019 — "Details of the Cloudflare outage on July 2, 2019" (698 points)news.ycombinator.com
- TechCrunch, Jul 2, 2019techcrunch.com



