An engineer deleted GitLab's production database. 300 gigabytes.
January 31, 2017, 23:27 UTC: a GitLab engineer, fighting a broken replica at the end of a long night, removes the PostgreSQL data directory on db1 instead of db2.
January 31, 2017, 23:27 UTC: a GitLab engineer, fighting a broken replica at the end of a long night, removes the PostgreSQL data directory on db1 instead of db2. db1 is the primary. About 300 GB of GitLab.com's database is gone in a second or two, and of the five backup and replication mechanisms, none is working. Postmortem: the timeline from the spam spike to the wrong hostname, why pg_basebackup looked stuck, why pg_dump had been failing silently (9.2 binaries on a 9.6 database, failure e-mails bounced by DMARC), the 18-hour restore from a 6-hour-old staging snapshot streamed live on YouTube, and who really gets the blame. Verdict on the response: SHIP IT.
Read the written edition (English) ↗
What this video covers
- Jan 31, 2017: rm -Rvf on the primary's data directory; ~300 GB removed, 4.5 GB left
- 5 of 5 backups fail: empty S3 bucket (pg_dump version mismatch), no Azure snapshots on the DB, wiped replica, daily LVM copy without webhooks
- Feb 1, 18:00 UTC: GitLab.com back from a 6-hour-old manual snapshot; live doc, live stream, blameless postmortem with a fix list
Transcript
0:00 An engineer at GitLab runs rm -rf on the wrong database server, and three hundred gigabytes of GitLab dot com vanish in a second or two, about how long it takes to read a hostname. January 31st, 2017, 11:27 p.m. UTC. GitLab tweets that it accidentally deleted production data, opens its incident notes to the internet, and streams the recovery on YouTube, the number two live stream on the platform. Next day, in writing: out of five backup techniques,
0:25 none are working reliably. How it happens, why it is possible, and who actually gets the blame. This is The Daily Diff, postmortem. 5:20 p.m.: an engineer snapshots production to test a load balancer in staging. 7 p.m.: spam hammers the database, plus a job hard-deleting a GitLab employee a troll reported for abuse. 11 p.m.: the replica falls so far behind that the primary has already discarded
0:48 the log it needs; the only fix is to wipe the replica and copy the primary again. pg_basebackup hangs with no output. It is actually waiting, silently, for the primary; nobody knows that, and the runbook does not say. The engineer, who meant to sign off at eleven, decides the empty data directory is the problem and removes it. On db1. The primary. He notices a second or two later; of roughly three hundred gigabytes,
1:11 4.5 remain. The backups. One: pg_dump to S3, daily. The bucket is empty. The cron job runs on an app server with no database, so the package picks PostgreSQL 9.2 binaries for a 9.6 database, fails, and emails the failure, which bounces for missing DMARC. Two: Azure disk snapshots, enabled for the file servers, not the databases.
1:32 Three: the replica, wiped on purpose an hour ago. Four: the daily snapshot, 24 hours old, every webhook stripped out by the staging sync. Five: the manual snapshot from 5:20, for an unrelated test. That one wins. Restoring means copying the staging disk back to production over Azure's cheap storage at sixty megabits per second: eighteen hours. GitLab dot com is back February 1st at six p.m. UTC, six hours of data older.
1:58 git blame: two hostnames one character apart, and five backup systems nobody has ever restored from. Not the engineer. The postmortem, signed by the CEO, keeps him anonymous, colours the production prompt red, and gives data durability an owner, because until now it had none. Blast radius: eighteen hours down, six hours of data gone, roughly five thousand projects, five thousand comments,
2:18 seven hundred new users, and five thousand people watching a progress bar. Hacker News gives the live doc 1,162 points and quotes one line back at them: out of five backups, none. Verdict, postmortem: ship it, on the response. They run the incident in public, blame the process, and publish the fix list with issue numbers. Monday action: restore a backup. If you have never restored it, you do not have one.
2:42 Send me the incident you are still not allowed to talk about, in the comments, or at the daily diff dot dev. And that's the diff for today. I'm Niko from Axrisi. Merge responsibly.
Sources
- GitLab, "Postmortem of database outage of January 31" (Feb 10, 2017)about.gitlab.com
- GitLab, "GitLab.com database incident" (Feb 1, 2017)about.gitlab.com
- @gitlabstatus, "We accidentally deleted production data…"twitter.com
- @gitlabstatus, emergency maintenance noticetwitter.com
- Hacker News, "GitLab Database Incident – Live Report" (1,162 points, 598 comments)news.ycombinator.com
- Hacker News, the postmortem thread (377 points)news.ycombinator.com



