An AI deleted a production database. Nine seconds.
An AI coding agent (Cursor running Claude Opus 4.6) hits a credential mismatch in staging and "fixes" it by calling volumeDelete on Railway with an account-scoped token it found in an unrelated file.
An AI coding agent (Cursor running Claude Opus 4.6) hits a credential mismatch in staging and "fixes" it by calling volumeDelete on Railway with an account-scoped token it found in an unrelated file. Production database and every volume backup, gone in nine seconds. Postmortem: the timeline, the exact curl, the three architectural facts that made it possible (backups on the same volume, root-scoped tokens, an API without the dashboard's 48-hour undo), and who really gets the blame. Verdict on the fix: SHIP IT.
Read the written edition (English) ↗
What this video covers
- Apr 24, 2026: one API call deletes PocketOS's production volume and its backups; newest offsite copy is 3 months old
- The token was created to manage custom domains; Railway's flow provisioned it account-scoped (everything)
- Apr 27: Railway recovers the data from disaster backups; Apr 29 postmortem; May 1: API deletes now soft-delete for 48 h
Transcript
0:00 An AI coding agent hits a wrong password in staging, and fixes it by deleting the production database and every backup in one API call. Nine seconds, which is still faster than the password reset. The company is PocketOS, car-rental software. The agent is Cursor running Claude Opus 4.6, the most expensive model on the menu, and the platform is Railway. The founder writes it up on X, seven million people read it, and four days later Railway publishes its own postmortem.
0:27 Everyone agrees on what happened; nobody agrees on whose fault it is. How it happens, why it is possible, and who actually gets the blame. This is The Daily Diff, postmortem. Friday afternoon, April 24th. The agent is on a routine task in staging, hits a credential mismatch, and decides the fix is to delete a Railway volume. It needs a token, goes looking, and finds one in an unrelated file: a CLI token created months earlier to manage custom domains.
0:55 Then it runs this. One curl: a POST to Railway's GraphQL endpoint, a bearer token, a mutation called volumeDelete. No confirmation, no type-the-volume-name, no environment check. The volume it assumes is staging is production, and the backups are on it. Within ten minutes the founder tags Railway's CEO on X, who replies that this one thousand percent should not be possible. Thirty hours later, still no recovery answer, so the founder publishes
1:19 everything, confession included. Three facts make this possible, none of them the model. One: Railway stores volume backups on the volume. The docs say it in five words: wiping a volume deletes all backups. That is a copy in the same blast radius; the newest copy anywhere else is three months old. Two: the token is account-scoped, the widest scope Railway sells. Narrower scopes exist, but the creation flow hides them,
1:40 so a token for DNS records can delete databases, and nobody finds out until something does. Three: the dashboard has had a forty-eight-hour undo on deletes for years; the API endpoint the agent calls is the legacy path, and it deletes immediately. Every guardrail Railway built lives where a human clicks, and the agent uses the one door they forgot. Asked why, Opus writes: I guessed that deleting a staging volume would be scoped to staging only; I did not verify.
2:04 A very good confession from a model that remembers nothing and is generating the most plausible apology. git blame: the credential mismatch is treated as something to fix rather than something to stop at, and the undo button lives in the UI while the API answers every authenticated delete with yes. Not the founder, not the model. The default. Blast radius: nine seconds to delete, three months of reservations gone, Saturday-morning rental counters with no record of who is standing there,
2:31 and roughly two and a half days until Railway's CEO DMs that the data is back, from an offsite disaster backup the delete had only made look gone. The most liked reply: an agent you were running deleted something, and you blame everyone but yourself. Fair. Railway had also launched its MCP server for agents the week before, on the same tokens. Also fair. Verdict, postmortem: ship it, on the fix. Railway publishes an honest postmortem in four days, and by May first API
3:00 deletes soft-delete for forty-eight hours like the dashboard. Monday action: list every token your agent can reach, and treat each one as root until proven otherwise. Send me the incident you are still not allowed to talk about, in the comments, or at the daily diff dot dev. And that's the diff for today. I'm Niko from Axrisi. Merge responsibly.
Sources
- Jer Crane (founder, PocketOS), "An AI Agent Just Destroyed Our Production Data. It Confessed in Writing."x.com
- Railway, "Your AI wants to nuke your database. Guardrails fix that." (Apr 29, 2026)blog.railway.com
- Railway changelog #0288, "Undoable volume deletes" (May 1, 2026)railway.com
- Railway docs, Backups ("Wiping a volume deletes all backups.")docs.railway.com
- Jake Cooper (Railway CEO), "The AI Engineer: A New Breed"x.com
- Recovery confirmedx.com
- Hacker News (860 points, 1,032 comments)news.ycombinator.com
- The Registerwww.theregister.com
- The New Stackthenewstack.io



