Armin Ronacher left GPT-6 Astra alone for 35 hours.
Armin Ronacher (Flask, Jinja) gave GPT-6 Astra one prompt and left it alone for 35 hours: about a billion tokens, about $1,200, 79 commits, 75,000 lines — and "absolutely nothing of value".
Armin Ronacher (Flask, Jinja) gave GPT-6 Astra one prompt and left it alone for 35 hours: about a billion tokens, about $1,200, 79 commits, 75,000 lines — and "absolutely nothing of value". The code it recovered is the story: C files patched by Python string-splicing heredocs, Python spawning Node spawning PowerShell, unit tests with the whitespace removed because it is 10% cheaper in tokens. Two measurements back him up: Earendil's SlopCodeBench numbers (agent code about twice as verbose and eroded as human repos, 0% strict pass rate) and Quesma's $1,500 benchmark of RTK (79k stars, "89% tokens saved" on its own counter, +17% per task with DeepSeek). Also: Cognition's SWE-2 (Kimi K3 in a trench coat, 92.8 on Terminal-Bench 2.1 vs 27.3 on Terminal-Bench 4), OpenAI's Agents API and paused $200 Pro sign-ups, and the Friday Hacker News asked for less AI news. Verdict: REVERT.
Read the written edition (English) ↗
What this video covers
- Armin Ronacher: 35 h of unattended GPT-6 Astra, ~$1,200, 79 commits, +75k lines, nothing shipped
- Earendil / SlopCodeBench: agent code ~2× as verbose and eroded as human repos; strict pass rate 0%
- Quesma: RTK reports 89% tokens saved, but Terminal-Bench costs go up 17% with DeepSeek; a rewrite loop failed 339 times in a row
- Cognition SWE-2: post-trained from Kimi K3, matches Fable 5.1 on FrontierCode at a third of the price; TB 2.1 92.8 vs TB 4 27.3; Devin Voice
- OpenAI Agents API (the Codex harness as a service, US-only, no ZDR); $200 Pro sign-ups paused for Astra demand
Transcript
0:00 Armin Ronacher, the man who wrote Flask, gave GPT-6 Astra a single prompt and left it alone for thirty-five hours. It came back with seventy-five thousand lines of code and, in his words, absolutely nothing of value, which makes it the most realistic software factory ever built. He posted the write-up on Wednesday. On Thursday, Cognition shipped SWE-2, and OpenAI shipped an Agents API and paused sign-ups for its two-hundred-dollar plan,
0:24 because Astra ate the servers. And on Friday the write-up reached the front page of Hacker News, a few rows below a post asking Hacker News to please stop posting about AI. In this video: what a day and a half of unsupervised Astra produces, two measurements that turn slop into a number, why a tool with seventy-nine thousand stars makes your bill bigger, and how Hacker News tried to filter out AI, using AI. It's Friday, September 11th, and this is The Daily Diff.
0:53 The factory. One prompt: build a Python with virtual threads and lexical scoping. Astra kept its own notes, spun up its own subagents, and ran until Armin pulled the plug, a day and a half and about twelve hundred dollars later. Seventy-nine commits, so fifteen and a half dollars per commit, which is roughly what a human charges, except the human eventually stops. The code is the story. Instead of the edit tool, Astra patches C files by piping Python heredocs that read the file, string-replace a function in, and write it back.
1:20 To run one Windows test it wrote Python that spawned Node that spawned PowerShell, a call stack with a passport. The unit tests it committed have no whitespace at all, because without indentation they're ten percent cheaper in tokens. And the task list drifted from one, two, three to task 8b2c2b2b checkpoint one, which is how you know the intern has been alone too long. Armin's theory: the model is rewarded for finishing long tasks and barely punished for ugly code, so token-golfed tool calls leak into the codebase,
1:48 and the fewer humans look, the less it matters. He titled that section It's AGI If You Don't Look. Hacker News added the economics: more code means more tokens to maintain it, which makes slop not a bug but a subscription. Can you measure slop? Armin's own company tried this week. Earendil ran the SlopCodeBench numbers: agent code is about twice as verbose and twice as eroded as the human repos it's compared against,
2:10 and when the benchmark wipes the agent's memory between checkpoints, the strict pass rate for every model they tested is zero. The most reliable slop metric they found was the raw line count, which stops working the moment anyone optimizes for it, so please don't tell the models. Then the wallet. RTK has seventy-nine thousand GitHub stars and YouTube thumbnails promising to halve your Claude Code bill; it compresses terminal output before the agent
2:34 reads it. Quesma ran it on Terminal-Bench for fifteen hundred dollars of tokens. With Fable the bill fell five percent, almost all from one task; with DeepSeek the average task got seventeen percent more expensive. Meanwhile RTK's own counter reported eighty-nine percent saved, because it counts bytes removed, not extra turns caused: a smart meter that bills the neighbour. And one version bug rewrote find into a command that failed, three hundred thirty-nine times in a row, at nine times the price.
3:01 The tool that kills tokens has a loop. The flood itself. Cognition's SWE-2 is Kimi K3, the open Chinese model, post-trained until it matches Fable 5.1 on Cognition's own benchmark at a third of the price; Hacker News: why are all American AI models basically Kimi in a trench coat. On Terminal-Bench 2.1 it tops the entire table; on Terminal-Bench 4, the one nobody has memorised yet, it scores twenty-seven to Fable's fifty-six,
3:28 so on the slide-deck benchmarks the best column is the exam everyone already passed. Cognition also announced Devin Voice, your favourite AI software engineer just got a landline, because the one thing agentic coding was missing was hold music. OpenAI, same day: the Agents API, which is the Codex harness sold as a service. OpenAI runs the sandbox, US data residency only, no zero-data-retention, so the agent remembers your codebase exactly as well as ChatGPT does. Then OpenAI paused sign-ups for the two-hundred-dollar Pro plan,
3:57 because Astra demand is, quote, unprecedented: the first product with a waitlist to pay more. Armin burned a full reset on his factory. He is the demand. Which brings us to Friday's Ask HN: can we please limit the AI news flood, six hundred fifty-six points, because, quote, every time someone at OpenAI or Anthropic farts, there's a top-ten post.
4:17 The replies were filters: hcker dot news removes AI stories with a BERT classifier trained on eight thousand examples, AI to delete AI, grass-fed; and a uBlock regex matching A-I case-insensitively also deletes email, daily, main, domain and train, so the regex, as always, is the villain. Same afternoon, four hundred fifteen comments argued about a Claude support page saying Claude is eighteen-plus.
4:38 The page is dated May. Nothing changed. Even the AI news that isn't news is news, which, to be fair, is also my business model. If you'd rather read this than hear me say it, the diff lands in your inbox every morning — free at the daily diff dot dev, link below. It is, unavoidably, AI news. So — today's verdict on the thirty-five-hour software factory: revert. Seventy-nine commits nobody read is not a codebase, it's a liability with a git
5:04 log; Astra is impressive, and the factory is the part to roll back. And that's the diff for today. I'm Niko from Axrisi. Merge responsibly.
Sources
- Armin Ronacher — Astra for Coding: Why Are We Doing This Again?lucumr.pocoo.org
- HN: Astra for Codingnews.ycombinator.com
- Earendil — Measuring the sloppiness of codeearendil.com
- HN: Measuring the sloppiness of codenews.ycombinator.com
- Quesma — RTK reports huge token savings, but our cost benchmarks disagreequesma.com
- HN: RTKnews.ycombinator.com
- RTK (Rust Token Killer)github.com
- Cognition — Introducing SWE-2cognition.com
- HN: Cognition SWE-2news.ycombinator.com
- OpenAI — Agents APIdevelopers.openai.com
- HN: OpenAI Agents APInews.ycombinator.com
- TechCrunch — OpenAI puts Pro subscriptions on hold due to Astra demandtechcrunch.com
- Ask HN: Can we please limit the AI news flood?news.ycombinator.com
- HN: Claude is only available to people over 18 yearsnews.ycombinator.com



