The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.
I have not been doing increasingly complex things since Opus 4.6 when models got really good.
My work at my job has stayed the same. But the model quality has varied.
They definitely tune the models in production after launch, if not only to share load during high traffic times. It’s not a crazy conspiracy that the same model can be stupider at different times.
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about.
Anecdata: I've been running a long-lived claude code session with Opus 4.6 for the last few days. Yesterday, almost right after the Sonnet 5.5 announcement, codex starting asking for permission to run things a lot more often
The quality of the output/work seems the same, but the speed at which is gets stuff done is a lot slower, because it's asking for permission so much more
I don't have any numbers/stats, just my impression. However, I imagine that if Anthropic could make the models ask for permission more often, it could be an interesting way to throttle access, without degrading quality of the output
I wonder if more organizations approving the model on a fast-tracked basis means Anthropic is straining for more compute and thus sheds a tiny bit to handle the increased demand, especially at peak times.
It seems as if this is based on demand. Whenever a new model is released, I'm guessing tens of thousands of us switch over to try the latest and greatest, which overloads the servers, leading to nerfing. It's 100% dishonest, but they realized they would lose users a lot quicker if they were honest and just said "our models are overloaded, come back later".
After Fable launch I switched over to Codex and it was simply amazing, with frequent usage resets that seemed never ending. They clearly had more compute than they knew what to do with. Post Astra, Codex has gotten dumb again across all models, increased usage for no real reason, and no resets.
I'm guessing Opus 5.5 will take the heat off Codex for a bit, leading to better performance. So I guess I stick around here instead of switching again?
(which has led me to believe that's a good approximation for hedonic adaptation, I've seen tons of attempts at demonstrating nerfing via benches, none persist)
I wonder if API is affected by this issue, especially Claude on public clouds? Would that means the subsidized rate just means they use cheaper quantized models and it's not comparable to API spending.
I've always used Enterprise per-token billing for Claude Code and I've never understood these nerf complaints. I've never noticed any slow downs at certain times of day, or a gradual decline in quality.
There’s probably contractual guarantees in the enterprise plans. My understanding of the subscriptions is they can swap the models out if any of them is getting too heavily loaded for a period of time
All this dishonesty and shadiness is part of why open models feel inevitable. Even if the total cost of ownership is higher (debatable; seems that way at small scales, but likely not as you grow), I'd rather have intelligence controlled by me that works for me.
The current period is as pro-customer as we're ever going to get, with cash still flying around and neither OpenAI nor Anthropic on the public market, and people are already forced into this sort of business to keep them true to their word.
Theory (Conjecture? Hypothesis?): What we notice as "model nerfing" is the company diverting compute to training/running new unreleased models..
Remember that some people get access to the next flagship version long before us peasants do. I recall seeing the mention of "Astra" more than a month before it was officially announced
Has there ever been any measurement of this, of any sort? Honest question. I frequently see a plural of anecdotes to that effect, but I've not seen a concrete statement of fact or measurement that could be scrutinized or tested in any way.
If so, please share. This should be measurable, and I'm glad this project is measuring it.
Answers in the form of additional anecdotes, stated with even greater passion but still lacking a statement that could be tested and falsified, would validate my exact concern.
The claim was that there's an experience of a model losing power. Your claim amounts to "No, you are not experiencing what you say". That's quite a claim for you to make with no data and no argument.
Yeh it's absurd that people claim this all the time. It's some crazy conspiracy theory and when you ask for examples nothing ever shows up.
It would be economical suicide from anthropic and OpenAI to actually need models intentionally.
But hey I guess it's hard with technology that truly seems like magic.
People say if you'd bring electricity to the middle ages you'd be called a witch and burned. The same is happening to the model labs here because they are bringing tech that the world isn't ready for yet.
I was just wondering if, like certain processors, bugs get fixed and the speed goes down. Like, they find it's doing things it shouldn't, restrict it, and harm the throughput.
I made a graphic to explain why people feel like the models get nerfed:
https://x.com/thesilenceturns/status/2103551351825543610
The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.
My work at my job has stayed the same. But the model quality has varied.
They definitely tune the models in production after launch, if not only to share load during high traffic times. It’s not a crazy conspiracy that the same model can be stupider at different times.
https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about.
The quality of the output/work seems the same, but the speed at which is gets stuff done is a lot slower, because it's asking for permission so much more
I don't have any numbers/stats, just my impression. However, I imagine that if Anthropic could make the models ask for permission more often, it could be an interesting way to throttle access, without degrading quality of the output
After Fable launch I switched over to Codex and it was simply amazing, with frequent usage resets that seemed never ending. They clearly had more compute than they knew what to do with. Post Astra, Codex has gotten dumb again across all models, increased usage for no real reason, and no resets.
I'm guessing Opus 5.5 will take the heat off Codex for a bit, leading to better performance. So I guess I stick around here instead of switching again?
(which has led me to believe that's a good approximation for hedonic adaptation, I've seen tons of attempts at demonstrating nerfing via benches, none persist)
The current period is as pro-customer as we're ever going to get, with cash still flying around and neither OpenAI nor Anthropic on the public market, and people are already forced into this sort of business to keep them true to their word.
Open models are the endgame.
Remember that some people get access to the next flagship version long before us peasants do. I recall seeing the mention of "Astra" more than a month before it was officially announced
If so, please share. This should be measurable, and I'm glad this project is measuring it.
Answers in the form of additional anecdotes, stated with even greater passion but still lacking a statement that could be tested and falsified, would validate my exact concern.
It would be economical suicide from anthropic and OpenAI to actually need models intentionally.
But hey I guess it's hard with technology that truly seems like magic. People say if you'd bring electricity to the middle ages you'd be called a witch and burned. The same is happening to the model labs here because they are bringing tech that the world isn't ready for yet.