2 ms·
I seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed m
by alexjplant 5d ago
I seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed massive peaks and troughs in performance week to week (albeit with Opus, not Fable).
I wonder what their official explanation for this behavior is.
- QwenGlazer9000 5d agoLast time they were called out, it was a regression in Claude code itself. At least that's their explanation. Either way, it wasn't a good look for "vibecoding" but it got brushed over.
- Wowfunhappy 5d agoWhen something is new, its capabilities feel incredible. Over time, those same capabilities become mundane, and you start to notice the flaws. (Now, if TFA is actually measuring reasoning tokens, that's quite different! It's not entirely obvious to me how he is measuring.)
- chrsw 5d agoI don’t think that’s what’s going on. I notice flaws on day one of model releases. But I also notice improvements if the model is truly more advanced than what I’m used to. Then over time the same questions or tasks return worse results. What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?
- Wowfunhappy 5d ago> What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage? ...I mean, if they were actually doing this despite saying that they don't—promising one product and delivering something else—I think that would be fraud, no? And, maybe it's one thing to secretly defraud normies like us (although class action lawsuits do exist), but I don't think major enterprises or the US military would take too kindly to it.
- pixl97 5d agoAre you telling me that companies might defraud people for millions and billions of dollars and pay fines that are 1000% less than their profits?" My goodness, you must live on a hell planet. Sorry there for the smarminess but fraud is just a standard business practice these days and fines are the cost of doing business. And I really am all for someone suing these companies forcing discovery so we can see how the sausage is made and how many eyeballs are in it.
- mobelkh 5d agois it? it's still the same model, they can claim the quantization down to q4 still retains 98% of the performance therefore it's fine. nothing on the fine print tells you what the weights are, you're just getting Fable 5, whatever that is
- himata4113 5d agoThey are deploying optimizations weekly (if not daily) with various AB tests. They don't manipulate model performance, but they do actively perform tests.
- espeed 5d agoThey did. More than once... Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude https://www.wired.com/story/anthropic-responds-to-backlash-on-claudes-secret-sabotage-on-ai-research/ https://www.wired.com/story/anthropic-responds-to-backlash-o... But it's still happening: https://github.com/anthropics/claude-code/issues/81759 https://github.com/anthropics/claude-code/issues/81759
- mirashii 5d agoAnd here's another great example of how a bunch of people who don't know what's going on throw noise into the system. That post is simply confused: the 1m opus calls are the auto-mode classifier, actual agent calls are still in Fable.
- espeed 5d agoLook at the usage. Fable wasn't being consumed.
- pixl97 5d ago>bunch of people who don't know what's going on Do you know why nobody outside the companies knows what's going on? Because they sell a black box with magic inside while steadfastly refusing to tell you if they are pushing buttons on said box while it is running. Can you imagine how much fraud would exist in the gambling industry if the gambling commission didn't exist at all? Everytime an industry is unregulated and has high costs of entry the entities in the industry abuse their customers. The incentives are much too high for them not to.
- bearjaws 5d agoYou're right to push back, and one honest caveat -- they could just be lying.