4 ms·
I think they neutered Fable. When it first came out it was indeed revolutionary. But what we have today, is not what we had before the ban.
by rubicon33 2mo ago
I think they neutered Fable. When it first came out it was indeed revolutionary. But what we have today, is not what we had before the ban.
- swader999 2mo agoNoticed that too. I wonder if these things just degrade over time, perhaps with the way it writes memories about my project as it goes
- Espressosaurus 2mo agoI’ve observed the degradation, but I suspect what’s happening is they’re tuning it for lower inference costs. Maybe turning down the amount of thinking, maybe quantizing, maybe something else. It seems like there’s a week by week and sometimes day by day change in performance when on a subscription plan using their harnesses.
- conception 2mo agohttps://marginlab.ai/trackers/claude-code/ https://marginlab.ai/trackers/claude-code/ their tracker generally shows that isn’t the case. The only times I’ve seen it drop is something broken and just before fable launched.
- nerdsniper 2mo agoI mean they could just be routing known benchmark questions (which all of SWEBench are) to a full-performance variant.
- svnt 2mo agoIs this using the api or using a subscription, though? The incentives are different for each, and it isn't the least bit unexpected that they would maintain API access quality while 'optimizing' the subscription experience to improve their margins (or losses) It seems to do really this you would need to crowdsource it -- users individually give the lab access to a body of subscriptions normally used by average people, and the lab occasionally runs some masked version of the task through on diverse accounts.
- 8note 2mo agoi thought i had noticed a degradation, but it turned out claude code had swapped itself back to opus. might be the case for you as well
- Gareth321 2mo agoI feel like a conspiracy theorist but it feels like every new model release from both Anthropic and OpenAI has 1-2 weeks of fantastic performance then a gradual (and sometimes not so gradual) decline in intelligence. It's like they're quantising the model in the background to optimise for available compute/RAM. Which, if I put my MBA hat on, would make perfect sense. Why not halve the required RAM for "90%" of the performance? They get fantastic benchmarks and glowing reviews on release, then slowly squeeze more performance out of the model. By the time the next model is ready for release, the jump feels quite large again.
- xnorswap 2mo agoAnd yet: https://marginlab.ai/trackers/claude-code-historical-performance/ https://marginlab.ai/trackers/claude-code-historical-perform... There's clearly random variation, but it also shows each model release is just genuinely better. ( With the exception of a heavily degraded week of opus 4.7, which was acknowledged as a problem at the time. ) There's a psychology of getting used to models after being wowed by the new performance. It sets in as your new baseline expectations, and then when it doesn't deliver, it's felt more acutely. When it does deliver, it's just meeting expectations. Then a new better model comes along and it's a step up again, another wow moment for a week or two until expectations adjust to meet the new baseline.
- jkman 2mo agoRemember https://en.wikipedia.org/wiki/Volkswagen_emissions_scandal https://en.wikipedia.org/wiki/Volkswagen_emissions_scandal? It's completely believable that benchmark-resembling requests are routes in a favorable manner.
- svnt 2mo agoI replied to the user above that referenced marginlab, but I believe marginlab uses the API. It is possible (arguably likely, in MBA-land) that the API and subscription accounts hit different sub-models. Even if they use a subscription account, surely Anthropic can tell which one it is.
- ACCount37 2mo ago
- jbird99 2mo agoAt this point I don't even bother with it. Constantly falls back to Opus anyway, so I may as well save myself some time.