3 ms·
I feel like a conspiracy theorist but it feels like every new model release from both Anthropic and OpenAI has 1-2 weeks of fantastic performance then a gradual
by Gareth321 3mo ago
I feel like a conspiracy theorist but it feels like every new model release from both Anthropic and OpenAI has 1-2 weeks of fantastic performance then a gradual (and sometimes not so gradual) decline in intelligence. It's like they're quantising the model in the background to optimise for available compute/RAM. Which, if I put my MBA hat on, would make perfect sense. Why not halve the required RAM for "90%" of the performance? They get fantastic benchmarks and glowing reviews on release, then slowly squeeze more performance out of the model. By the time the next model is ready for release, the jump feels quite large again.
- xnorswap 3mo agoAnd yet: https://marginlab.ai/trackers/claude-code-historical-performance/ https://marginlab.ai/trackers/claude-code-historical-perform... There's clearly random variation, but it also shows each model release is just genuinely better. ( With the exception of a heavily degraded week of opus 4.7, which was acknowledged as a problem at the time. ) There's a psychology of getting used to models after being wowed by the new performance. It sets in as your new baseline expectations, and then when it doesn't deliver, it's felt more acutely. When it does deliver, it's just meeting expectations. Then a new better model comes along and it's a step up again, another wow moment for a week or two until expectations adjust to meet the new baseline.
- jkman 3mo agoRemember https://en.wikipedia.org/wiki/Volkswagen_emissions_scandal https://en.wikipedia.org/wiki/Volkswagen_emissions_scandal? It's completely believable that benchmark-resembling requests are routes in a favorable manner.
- svnt 3mo agoI replied to the user above that referenced marginlab, but I believe marginlab uses the API. It is possible (arguably likely, in MBA-land) that the API and subscription accounts hit different sub-models. Even if they use a subscription account, surely Anthropic can tell which one it is.
- ACCount37 3mo agoYou should. Feel like a conspiracy theorist when saying things like this. Users are not reliable or consistent model evaluators. Users adapt to models - the moment they get a model that performs better, their expectations rise, their tasks get harder and their prompts get shorter. "They made the model worse" is PEBKAC in 9 cases out of 10.
- svnt 3mo agoThere are few conspiracies where the vectors between "capitalist organization makes more money" and "user can't reliably distinguish tiers of product quality" overlap.
- ACCount37 3mo agoI wish the users weren't so fucking stupid with the "they made the model worse" stuff. Then that 1 out of 10 case where the model was actually made worse (whether intentionally or by mistake) would stand out instead of being swallowed by the noise floor.
- dmrivers 3mo agoThis is also my experience. I don't know if it's because of the quantization theory, or if it's just me getting used to a certain level of coding performance and gradually less tolerant of the mistakes it makes more over time.