Where the word comes from
In online games, developers regularly rebalance their games with patches. When a weapon or character gets weaker, players say it was “nerfed”, after the soft foam darts of Nerf toys. The term moved to AI when users started to notice that a model they relied on behaved differently weeks after launch, without any announcement.
What is documented
A few cases are on the record because the providers explained them or researchers measured them:
- Anthropic, September 2025. Three infrastructure bugs degraded Claude’s answers between 5 August and 18 September 2025. At the worst point, 16 % of Sonnet 4 requests were affected. Anthropic published a detailed postmortem.
- Anthropic, April 2026. The default reasoning effort in Claude Code was lowered from high to medium for about a month to reduce waiting times. Anthropic later called it “the wrong tradeoff”.
- OpenAI, April 2025. An update made GPT-4o overly flattering. OpenAI rolled it back within days and explained what went wrong.
- Stanford and UC Berkeley, 2023. Researchers measured clear changes in GPT-4’s behaviour between the March and June 2023 versions of the same model name.
What is suspected but rarely confirmed
Users and some independent trackers suspect other mechanisms: stronger quantization (running a model with lower numeric precision to save compute), routing part of the traffic to smaller models, or quality that depends on the time of day and server load. These are plausible engineering trade-offs, but providers rarely say whether, when or for whom they apply. That uncertainty is the core problem.
Why it is hard to prove
Language models answer differently every time, so a single bad answer proves nothing. Expectations also change: a model that impressed in its first week can feel ordinary a month later even if nothing changed. Without published version information and changelogs, users cannot separate a real change from chance, and providers cannot easily show that nothing changed.
What would fix it
Transparency: pinned versions, a public changelog for how models are served, advance notice, quality measurements anyone can check and honest incident reports. That is what the open letter asks for. If you want to collect your own evidence, start with Document it yourself.
Frequently asked questions
What does “nerfing” mean for AI models?
The word comes from video games, where a patch “nerfs” a character or weapon by making it weaker. For AI models it describes a drop in quality after release while the model name and the price stay the same. Users notice shorter answers, ignored instructions or more mistakes on tasks that used to work.
Do AI providers nerf their models on purpose?
Sometimes a change is deliberate. In April 2026 Anthropic reported that it had lowered the default reasoning effort in Claude Code from high to medium for about a month to reduce waiting times, and later called that “the wrong tradeoff”. Other drops were caused by bugs, such as the three infrastructure issues Anthropic described in September 2025. Many other cases are suspected but not confirmed, because providers rarely disclose how a model is served.
Why would a model get worse after launch?
Documented causes include infrastructure bugs, changes to default settings such as reasoning effort, system prompt changes and model updates that were later rolled back. Commonly suspected but rarely confirmed causes are stronger quantization, routing some requests to cheaper models and limits that depend on server load. Randomness and changing expectations can also make a model feel worse when it is not.
How can I tell if a model was nerfed?
Save a set of real tasks with exact prompts and results on launch day. Repeat them regularly in fresh sessions, several times each, and note dates and response times. Compare with a pinned, dated model version through the API where possible. Our guide “Document it yourself” explains the steps.
What is a pinned model version?
A pinned or dated version is a model identifier that refers to one fixed snapshot, for example a name that ends in a date. It should behave the same for its documented lifetime. Aliases like “latest” can point to a different model at any time.
What does EU law say about changes to paid digital services?
Directive (EU) 2019/770 allows a provider to change a digital service during a subscription only if the contract allows it and customers are clearly informed. If a change impairs the service more than slightly, customers must be told in advance and may cancel within 30 days, unless the unchanged version stays available at no extra cost. This is general information, not legal advice.
What is Stop Nerfing asking for?
Pinned versions, a public changelog for how models are served, advance notice before changes that can lower quality, quality measurements that anyone can check, and honest incident reports. The open letter explains each point and lists its sources.
Who is behind Stop Nerfing?
Stop Nerfing is a non-commercial campaign run by Kindler Web Services in Unterhaching, Germany. It is not affiliated with any AI provider.