Answer

Do AI models get worse after launch, and how can you tell?

Sometimes, yes, but rarely in the way people assume. The best-documented cases are serving bugs and silent version changes: Anthropic published a postmortem of three 2025 infrastructure bugs that degraded Claude responses for weeks, and a 2023 Stanford and Berkeley study measured large swings in GPT-4 between March and June. Many other reports turn out to be noise, changed tools or changed prompts. The only reliable way to tell is to run a fixed set of tasks repeatedly and compare the results with error bars.

Published · Updated · Evidence-linked, not search-volume ranked.

Short answer

AI models can get worse after launch, but the documented cases are usually not a deliberate downgrade of the same model. In September 2025 Anthropic published a postmortem of three infrastructure bugs that intermittently degraded Claude responses between August and early September 2025: a routing error that at its worst hour hit 16% of Sonnet 4 requests, an output corruption bug, and a compiler bug in token selection. Anthropic says it never reduces model quality due to demand, time of day or server load. Separately, a 2023 study by Lingjiao Chen, Matei Zaharia and James Zou found that GPT-4 answered a prime-number task correctly 84% of the time in March 2023 but 51% in June 2023, showing that a service with the same name can change substantially. Other causes of a model feeling worse include a new version of the app or coding tool around it, different default settings, a longer or messier conversation, and ordinary randomness. You cannot tell which one applies from a few bad answers. Run the same fixed tasks on a schedule, pin every version you can, record accuracy and output length, and compare against an early baseline with error bars.

Why this question is current

Exact query-volume data was unavailable, so RepoRadar uses these as current demand and intent signals rather than a claimed volume ranking.

  • is claude getting worse · Google Suggest · US; English · checked 2026-10-01T22:27:13Z
    Observed completions: is claude getting worse, is claude getting worse reddit, is claude getting worse at writing, is claude getting worse recently, does claude get worse over time, is claude code getting worse, is claude ai getting worse, is claude opus getting worse, is claude sonnet getting worse, is claude usage getting worse. A formulation signal captured at this time, not a volume or ranking claim.
  • is opus 5.5 nerfed · Google Suggest · US; English · checked 2026-10-01T22:27:13Z
    Observed completions: is opus 5.5 nerfed, plus two unrelated general completions. A formulation signal captured at this time, not a volume or ranking claim.
  • opus 5.5 nerf · Google Suggest · US; English · checked 2026-10-01T22:27:13Z
    Observed completions: opus 5.5 nerfed. A formulation signal captured at this time, not a volume or ranking claim.
  • ai model nerfed · Google Suggest · US; English · checked 2026-10-01T22:27:13Z
    Observed completions: ai models nerfed, plus unrelated general completions. A formulation signal captured at this time, not a volume or ranking claim.
  • stories with more than 50 points, trailing 48 hours · Hacker News Algolia search_by_date · global English-language developer community · checked 2026-10-01T22:25:50Z
    Story 49901736, Livenerf: Has Opus 5.5 been nerfed yet? (github.com/ninjahawk/livenerf), created 2026-09-29T22:36Z, 909 points and 388 comments at check time, the second-highest story in the 48-hour sample. Interest signal, not search volume.

Who this helps

  • developers who suspect a model or coding agent got worse and need evidence before switching
  • teams running AI features in production that need to catch quality drift early
  • everyday users trying to decide whether a bad week is the model, the app, or the prompt

What nerfed means here

Nerfed is gamer slang for something made weaker in an update. Applied to AI, it is the suspicion that a provider quietly serves a cheaper or weaker version of a model some weeks after launch, once the launch reviews are written. People name possible mechanisms such as quantization, a smaller model behind the same name, less reasoning effort, or different routing.

That suspicion is a hypothesis, not a finding. It is worth separating three questions: did the output actually get worse, if so what changed, and was the change deliberate. Most public arguments skip straight to the third.

What has actually been documented

The clearest public case is Anthropic's September 2025 postmortem. It describes three overlapping infrastructure bugs from August to early September 2025. Some Sonnet 4 requests were sent to servers configured for a 1 million token context window; at the worst hour on August 31 that affected 16% of Sonnet 4 requests, and about 30% of Claude Code users in the period had at least one misrouted message. A second bug occasionally produced stray Thai or Chinese characters or obvious syntax errors. A third, a compiler bug in how the next token was chosen, affected Claude Haiku 3.5 and possibly other models.

Anthropic's own explanation of why it took weeks to fix is instructive: its evaluations did not catch what users reported, partly because Claude often recovers from isolated mistakes, and privacy controls limited engineers' access to the problem conversations. The company also stated that it never reduces model quality due to demand, time of day or server load. That is a statement from the vendor, not something outsiders can verify directly.

The second well-known data point is academic. In a 2023 paper, Chen, Zaharia and Zou compared the March and June 2023 versions of GPT-3.5 and GPT-4. GPT-4 fell from 84% to 51% on identifying prime numbers, while GPT-3.5 improved on the same task, and both had more formatting mistakes in generated code in June. The authors concluded that the behavior of the same named service can change substantially within months and needs continuous monitoring.

Other reasons a model can feel worse

This section is analysis rather than reported fact. Before blaming the model, check what else changed around it.

  • The tool changed. A new release of a chat app or coding agent can change the system prompt, tools and defaults. The livenerf project pins its Claude Code CLI version for exactly this reason, noting that a changed harness looks like a changed model.
  • The model name points somewhere new. Alias names are designed to move: Google documents that its latest aliases are hot-swapped with each release after a two-week email notice, while specific stable names usually do not change.
  • Settings changed. Lower reasoning effort or a smaller output limit produces shorter, shallower answers from the same model.
  • Your context changed. Long conversations, larger repositories and messier instructions make every model worse.
  • Expectations changed. Launch week tends to involve easy showcase tasks; later you hit the hard cases, and a few bad answers from a random process are easy to over-read.

How the livenerf project tests it

Livenerf, a public GitHub project that drew wide discussion this week, tries to answer the question for Claude Opus 5.5, which launched on September 22, 2026. It started a daily run on September 24 against a fixed panel of tasks with exact graders, using a pinned Claude Code version, and plans 30 days: days 1 to 10 as the baseline and two later 10-day windows. Its pre-registered rule only calls a change if the 99% interval excludes zero in two consecutive windows, the effect is at least 3 points, and a control arm does not move the same way.

Two details are worth copying. It tracks output tokens per sample as an early warning, on the theory that a model which starts thinking less shows it there before accuracy moves. And it treats launch week as a reference point rather than ground truth, since launch week can be the worst week. As of its last README update on September 27 it had 4 of 30 days and no results; the first possible call is around October 24. It measures Opus 5.5 through a Claude subscription in Claude Code, not the raw API, and its license is not yet chosen.

Run your own small drift check

You do not need a research project to stop guessing. A modest version fits in an afternoon and gives you evidence you can act on or report.

  • Pick 20 to 50 real tasks with answers you can check automatically, such as a failing test that must pass or a field that must be extracted exactly.
  • Pin what you can: a specific model ID rather than an alias, the client or agent version, the reasoning setting and the system prompt. Write them down.
  • Run every task several times, because one run of a random process proves little. Repeat on a schedule, for example weekly.
  • Record the pass rate and the output length for each run, and save the raw outputs.
  • Compare against your earliest runs with error bars. Evan Miller's paper Adding Error Bars to Evals explains how to measure the difference between two sets of runs. A drop inside the noise is not evidence.
  • If you find a real drop, report it with the evidence. Anthropic asks for reports through the /bug command in Claude Code or the thumbs down button, and says specific reports helped it isolate the 2025 bugs.

Limits of this answer

RepoRadar has not measured any model for drift. The documented cases above are historical and concern specific models in 2023 and 2025; they do not show that any current model, including Opus 5.5, has been degraded, and they do not show that one has not. Vendor statements about not reducing quality cannot be independently verified. Livenerf has no results yet and is a single-person project testing one model in one tool.

A useful next action

If you depend on a model for real work, set up the small drift check this week, while your current results are your baseline. Next time it feels worse, run the panel before switching providers or rewriting prompts. For a fuller method for judging agents before you trust them, see the RepoRadar answer on evaluating an AI agent before production.

Sources checked

  • Anthropic Engineering: A postmortem of three recent issues (September 17, 2025) ↗ checked · vendor engineering report, global

    Primary source for the three August to September 2025 infrastructure bugs, the 16% worst-hour and roughly 30% Claude Code user figures, why detection was slow, the statement that quality is never reduced due to demand, time of day or server load, and the request to report via /bug or thumbs down.

  • arXiv: How is ChatGPT's behavior changing over time? (Chen, Zaharia, Zou, 2023) ↗ checked · research paper, global

    Measured March versus June 2023 behavior of GPT-3.5 and GPT-4, including GPT-4 falling from 84% to 51% on prime identification, and concluded the same named service can change substantially in a short time.

  • GitHub: ninjahawk/livenerf ↗ checked · public project repository, global

    Describes the 30-day Opus 5.5 drift benchmark, its start date, pinned Claude Code CLI, pre-registered decision rule, output-token secondary signal, progress of 4 of 30 days as of 2026-09-27, and that its license is not yet chosen.

  • arXiv: Adding Error Bars to Evals (Evan Miller, 2024) ↗ checked · research paper, global

    Statistical guidance for analyzing evaluation data and measuring differences between two models while minimizing noise.

  • Google AI for Developers: Gemini API models, version name patterns ↗ checked · official API documentation, global

    Documents that stable model names usually do not change while latest aliases are hot-swapped with each release after two weeks of email notice.

RepoRadar separates factual source claims from analysis. Recheck vendor docs before purchase, deployment, or policy decisions.