3 ms·
New release of fable and opus 5.5 is pending and Anthropic is reallocating resources. Degradation always happens in transition, it sucks. Opus 5.5 is being ser
by prodigycorp 10d ago
New release of fable and opus 5.5 is pending and Anthropic is reallocating resources. Degradation always happens in transition, it sucks.
Opus 5.5 is being served under opus 5 right now.
- w1296 10d agoEspecially with the frequent releases aka version bumps.
- deleted 10d ago[deleted]
- SequoiaHope 10d agoCan you elaborate on the mechanism of this degradation? If resources are not available I would expect a request to fail with a message about resources not available. Do they tweak back end model capabilities to maintain service in a degraded state?
- arcanemachiner 10d agoDollars to donuts, they are speculating, and not privy to inside information on the topic. However, I believe that runtime model quantization is possible with some publicly-available inference engines (e.g. vLLM), so its not beyond belief that the closed labs do quantize at runtime, either to allocate compute, or to nudge users towards a preferred model (e.g. make the incumbent model dumber to push people to use the latest-and-greatest model, or vice versa to ease the load on the latest model, which is typically larger than the old one).
- rybosworld 10d agoAn AI lab will never volunteer the information because it opens them up to lawsuits if they are purposely degrading service and not letting users know. They can limit how hard the model thinks for a given effort. Suddenly xhigh only thinks as hard as high did, and high shifts down to medium effort, and so on. They can also serve quantized models. And this has the benefit of practically not showing up in benchmarks at all even if the user experience is obviously degraded. The other major thing the labs do is silently drop the usage limits. This has become very noticeable for codex users who are suddenly burning through their weekly usage in a few hours.
- pixl97 10d agoYea, if you ever run your own models on a GPU there are a whole ton of different dials you can adjust that drastically affect compute use, memory use, and output token quality, and number of tokens held in memory. If anyone reading has a GPU it's worthwhile just messing with a smaller model for a bit to watch how the settings affect output.
- gslepak 10d ago> Opus 5.5 is being served under opus 5 right now. On what basis are you claiming this?
- pllbnk 10d agoIt shouldn’t be an excuse. They are selling a product and that product should always be within the quality range.
- user43928 10d agoAnd it's not. A conspiracy theory is what it is. I have no reason to doubt the claims of the employees at OpenAI and Anthropic who have told us personally multiple times, including here on HN, that they do not degrade the models in order to reduce load. As for the endlessly long analysis in the OP, it appears it's based on analyzing their random usage data rather than any fixed benchmark. I don't think it makes much sense.
- pertymcpert 10d agoWhy would reallocating resources make a single inference run worse in quality?