How to detect a Claude nerf (with real data)
People say Claude gets dumber a few weeks after every launch.
People say Claude gets dumber a few weeks after every launch. One developer put Claude Opus 5.5 on a 30-day clock from launch week, with frozen questions, a pinned Claude Code and public rules, and it shows why nobody could have known before, and why your token count is the thing to watch. Verdict: SHIP IT.
Read the written edition (English) ↗
What this video covers
- Claude gets dumber after launch: the theory meets a day-zero clock
- livenerf: a 30-day clock on Opus 5.5 from launch week
- How to catch a nerf: 78 sometimes-right questions
- What it can see: 7.5 points, and the token count moves first
- The blind spot: an Opus 5 swap and 8 wrong answer keys
Transcript
Claude gets dumber after launch: the theory meets a day-zero clock
0:00 Everyone knows Claude gets dumber a few weeks after launch. So one developer put Opus five point five on a clock from launch week, and the clock says nobody could have known yet. In this video, three questions. How do you catch a nerf? What can this meter see? And what does Anthropic say? And one line in the repo's credits made me laugh out loud.
0:20 I'll get to it at the end. It's Wednesday, September thirtieth, and this is The Daily Diff. Three stories today, and the big one is a GitHub repo called live nerf.
livenerf: a 30-day clock on Opus 5.5 from launch week
0:30 For months, people have said Anthropic quietly nerfs its models after launch. The usual theories are a squeezed model, a smaller model under the same name or simply less thinking. The repo calls every argument so far vibes versus vibes. Opus five point five shipped on September twenty-second, and a week ago I told you it topped the independent board while writing three times more words than the median model. Two and a half days in, a developer called ninja hawk started the clock,
0:59 once a day for thirty days, with logs nobody may edit. The Hacker News thread hit almost eight hundred points and split like a team chat. One commenter says nerfing models isn't real in the vast majority of reported cases. Another says people are just getting used to the new level of intelligence. And one Claude Code power user now dismisses the rate-this-session pop-up every time, because rating Claude seemed to make it worse. Peer review. So, question one, how do you catch a nerf?
How to catch a nerf: 78 sometimes-right questions
1:27 You start with about twenty-three hundred hard exam questions, and Opus gets ninety-three percent right on the first try, which is useless for a detector, since a question it always gets right can't drop. Ninety-seven percent were always right or always wrong, and the seventy-eight it only sometimes gets right became the panel. Same prompt, no tools, and exact-match grading with no AI judge, because the judge would drift too. It runs on a Max subscription through headless Claude Code,
1:55 since the API version was priced at about sixteen hundred dollars a month. And the Claude Code version is pinned, because, in the repo's words, a changed harness looks exactly like a changed model. The statistics come from a paper called Adding Error Bars to Evals, by Evan Miller at Anthropic. So Anthropic's own math is now auditing Anthropic, which is the academic version of quoting the terms of service back at customer support.
What it can see: 7.5 points, and the token count moves first
2:19 Question two, what can it see? Per ten-day window, a drop of about seven and a half points. To prove the rig works, the author turned Claude's effort down on purpose. Medium effort cut the output by about a quarter, and the score by only four points. Low effort cut the output by almost two thirds, for eight points. That's the part to steal. Less thinking shows up in the token count before it shows up in the answers.
2:43 So pin your model version, log output tokens per task, and keep a few frozen questions of your own. If your agent suddenly writes shorter, check your logs before you check Reddit. And here's the honest part.
The blind spot: an Opus 5 swap and 8 wrong answer keys
2:54 Quietly swapping in the older Opus five was too small a change for the rig to call, and the readme says so up front. The audit also found eight answer keys that look wrong, so the nerf detector found bugs in the benchmark before it found a nerf.
Anthropic: we never reduce model quality
3:08 Question three, what does Anthropic say? Last year, after a wave of complaints, Anthropic published a postmortem. It says, quote, we never reduce model quality due to demand, time of day or server load. The same post admits three bugs had degraded Claude, and almost a third of Claude Code users hit the wrong servers at least once. So the complaints were real and the cause was bugs, which is why the repo warns that launch week could be the worst week.
3:35 Six of thirty days are in. The first possible call lands around October twenty-fourth, so the honest answer is that nobody knows, including everyone who was sure. Speaking of numbers nobody publishes.
Data centers keep their numbers secret; Backblaze publishes
3:46 Lighthouse Reports and the Dutch paper Trouw found that the vast majority of European data centers keep their power and water use secret, although an EU rule has required reporting for three years. In the Netherlands, fewer than a quarter publish. One Microsoft data center alone uses about one percent of Dutch electricity. Meanwhile, Backblaze did the boring thing again. Its quarterly drive stats cover about three hundred fifty thousand hard drives, and the failure rate rose to one point seven three percent,
4:16 the highest in a while. Three Seagate models had zero failures. That's what a public baseline looks like, quarter after quarter, for thirteen years.
The credits line: Claude helped build the meter
4:25 Now, that credits line. The readme admits a lot of the repo was written with the help of Claude, which is the model being measured. So Claude helped build the meter that checks whether Claude got dumber. That's why the graders are plain functions. In the author's words, you shouldn't have to trust the author, human or otherwise. If you'd rather read this than hear me say it, the diff lands in your inbox
4:46 every morning, free at the daily diff dot dev, link below. So, today's verdict on live nerf.
Verdict: SHIP IT, and don't quote it before Oct 24
4:51 SHIP IT. I'd ship it because it's the first nerf argument I've seen with a timestamp, a public rulebook and a stated blind spot. Just don't quote it before October twenty-fourth. And that's the diff for today. I'm Niko from Axrisi. Merge responsibly.
Sources
- livenerfgithub.com
- Validationgithub.com
- Hacker Newsnews.ycombinator.com
- Anthropic postmortem (Sep 2025)www.anthropic.com
- Adding Error Bars to Evals (Evan Miller)arxiv.org
- Inspect (UK AI Security Institute)inspect.aisi.org.uk
- NL Timesnltimes.nl
- Hacker Newsnews.ycombinator.com
- Backblaze Drive Stats Q2 2026www.backblaze.com
- Hacker Newsnews.ycombinator.com



