3 ms·
Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better th
by talon8635 5d ago
Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?
For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.
I have no idea if that’s what’s happened, I completely pulled it out of my butt. And I have no idea is the actual frontier is stagnating.
- AmazingTurtle 5d ago> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one? Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out. Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.
- ramesh31 5d ago>Also I will likely save some money on subscriptions. Unlikely. The $200 Claude subscription allows for billions of tokens/month, and that kind of hardware will take years to amortize.
- torben-friis 5d agoThere could be gym logic at play. Hundreds signed up, 20 people actually exercising. Though it's probably more likely in the lower tiers.
- rybosworld 5d agoI wouldn't be so sure. The generosity of the subscription plans has declined GREATLY over the past 6 months or so. They are likely trending towards api pricing parity. In which case, having your own hardware makes sense if you can utilize it well.
- zeroonetwothree 5d agoLast time I estimated it was like 30 years to pay back. I doubt the hardware will even last that long.
- fragmede 5d agoLast time I estimated, it would only take 3 months to pay back because the 1TB Mac Mini running Qwen RSIingly developed ASI and made infinity dollars off of crypto and I got put in jail by the SEC. Where'd you get 30 years from? Show your work.
- wilj 5d agoI would like to subscribe to your newsletter.
- mike_d 5d agoI have 2 x ChatGPT Pro 20x, Claude Max 20x, and Kimi Vivace. It's about ~12 months payback for two units and the cable. The problem is they can't fit any frontier level open models.
- AtHeartEngineer 5d agoflash next is good, I've been running it for like 2 weeks now and it's pretty solid, hope you like it and it meets your needs. I still lean on Claude and codex a fair bit for harder stuff, but I'm rapidly moving towards 2x $20 plans instead of 2x $200 plans
- boardwaalk 5d agoI don’t know what people do with the open models but having tried a lot of them I just can’t make it make sense. they’re too dumb and it effectively makes them useless (to me). it’s probably worth being honest about the low ceiling here.
- cyanydeez 5d agoQwen3.8-Flash-Next seems pretty much auto pilot when I get it the right context. Perhaps reverse the question: Are your build/construct requirements just really counter-productive to how LLMs need to understand things? I've found constructing the code, writing the tests, adding the docs; then running through them gets most of the way there. I've also found that making a simple obvious edit is a useless endevour when the LLM is primed for the long context tasks. So, again, the question is reversed: are you over reliant on the LLM to do even stupid simple likes like editting a css variable?
- redanddead 5d agoServing compute is their main value prop Yet… even Altman called out Anthropic for serving dumbed down models. Shits weird man
- tsunamifury 5d agoEvery SOTA model I've used at launch uses deeper, longer inference then gradually turns down over time, until the next model comes out which seems to be trained on some new data, but mostly performance due to deeper longer inference for another period.
- briffle 5d agoI have not been attributing it so much to malice, just that all the major cloud vendors seem to be running at full capacity, and can't build new datacenters fast enough. I just kind of assumed that as they got busy training newer models, that they allocated less resources to handle the existing systems, because they aren't able to get more capacity right now.
- Denzel 5d agoI’m not sure why this point keeps coming up — if your service/product is so popular that it’s capacity-constrained, then the answer is to raise prices, not degrade service, because the demand should be inelastic.
- pixl97 5d agoThis really depends where the load shedding point is. A very small raise in prices may cause a very large loss in customers that you risk never getting back. For example if customers figure out that the Chinese models are just as good, they are gone because they are so much cheaper.
- bonoboTP 5d agoRaising prices also has second order effects, like consumer and business expectations around how widespread the tech can be. Valuations depend on it being reasonably affordable to roll out on a much more massive scale than today. If people get the impression that it seems too limited to very rich people (200 is affordable for a North American / Western European professional), the impression about the trajectory will change.
- fnordpiglet 5d agoThe Opus 4-6,4-8,5 arc is exactly this. As one person commented in here, opus 5 is a terrorist. This is undeniable. Opus 4-6 was awesome. 4-8 was worse behaviorally but produced better code. Fable seems to be following the same enshittification arc of other Anthropic models. Generally OpenAI seems to be taking the opposite approach with an increasing improvement over time. As sad as I feel to say this, open ai seems to have the right strategy. Making your product worse over time rarely plays well with customers. At this point it feel often hard to justify using Anthropic for anything. I generally like Anthropic better as a company and they really had the initiative and advantage and customer good will, then proceeded to squander it faster than a cigarette company or the Sacklers could have.
- fragmede 5d ago> Making your product worse over time rarely plays well with customers. On the other hand, New Coke was a resounding success. Well, it, itself wasn't, but in the aftermath, Coke outsold Pepsi 2:1.
- talon8635 5d agoI’m so behind on this topic but I find it interesting how quickly things change. I feel like just yesterday I way hearing how anthropic is far and away better than OAI, and now this. I have no way to judge myself. I don’t even use them. But it’s interesting to follow by just reading stories and comments
- progval 5d agoThis sounds similar to rumors about how SSD companies work. First they would design a new drive with better performance that everyone uses to benchmark against other models; then slowly change its parts to worse ones, either because they are cheaper, the originals are no longer available, or whatever reason
- nxc18 5d agoThere must be some benefit if all the providers are doing it independently. GPT5.6-Sol on Max thinking just became regarded as of a few days ago. The boosters will tell me it’s my fault for using such an old, cheap out-of-date low quality near useless wish.com model (that was SOTA and better than human coders one month ago). The cycle repeats.
- deleted 5d ago[deleted]
- deleted 5d ago[deleted]
- talon8635 5d agoAgain, I’m out of my element here, but isn’t the entire industry dependent on “new better releases frequently”? If so, and if no one has made any meaningful breakthrough, might they all pursue this kind of deception just to stay afloat/“competitive”/relevant? Thanks for your insight
- kleiba2 5d ago> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one? Not unless your competitors do the same, or else you will only be perceived as falling behind others.
- talon8635 5d agoYes that makes sense. In my hypothetical, the industry frontier is stagnating, meaning no one is making big breakthroughs, so they all resort to this. If one lab makes a breakthrough, can the other labs just distill to bear parity anyways and then set a new baseline industry wide. I’m quite ignorant on this topic, so if any of this sounds moronic, forgive me
- wmf 5d ago...releasing a new model that’s marginally if at all better than the original... This isn't what we see in benchmarks.
- rfgplk 5d agoYep, that's what they've been doing for a long while now. Also the amount of tokens you get per sub varies drastically from month to month. Needs to be regulated.
- ekjhgkejhgk 5d ago> For an industry that’s stagnant in progress Yes, the AI technology is known primarily for how stagant it is.
- talon8635 5d agoYes, I freely admitted I was entertaining a pure hypothetical I pulled out of my butt. I have no idea, just had a thought and put it out there
- whatever1 5d agoYou can serve Fable from a cloud vendor (like AWS, Azure). They have frozen versions of the models, so likely this should not be an issue? I would do a test to verify my suspicions.
- talon8635 5d agoSounds like a good smoke test. I’m actually so far removed from this tech that I couldn’t run such a test myself lol
- Aurornis 5d ago> to create a perceived improvement when in reality there isn’t really one? This wouldn't explain progress on benchmarks (including closed sets), or the fact that newer models are providing solutions to major math problems that older models cannot.
- pixl97 5d agoFar more likely it's about reducing costs.
- well_ackshually 5d ago* Release new model that scores an arbitrary 100 on a benchmark * Get everyone to talk about you as the first model to ever score 100 on the 100benchmark. * Tune it down over time so that you end up only scoring 75 on the benchmark and people get used to it, gaslight them into thinking it never changed or that it's just a harness problem, they can't run the old version locally anyways to verify. This also cuts your costs in half. Your gross margin on API calls goes from 70% to 150%. * Release new model that scores 120 on the benchmark and advertise it as 50% better than the current model, while it's only in practice a minor increment. Everyone praises it as the second coming of Jesus Christ. * Get everyone to talk about you as the first model to ever score 120 on the 100benchmark. Bis repetitae.
- scrollop 5d agoWhy can't the models be benchmarked again after a few weeks/months to confirm this (likely true) theory? I imagine some people have their own personal in depth benchmarks they could do this for.
- Aurornis 5d ago> * Tune it down over time so that you end up only scoring 75 on the benchmark Where? I see so many accusations of this happening and it's so easy to check, but nobody ever proves it.
- mobelkh 5d agobut there is a gap between benchmarks and user feel. Opus 5 came out with better benchmark results than Fable, but it really did not feel better to use at all.
- holoduke 5d agoNah its because they cache and preprocess requests by dumb models and send them too often to another dumb models instead of the top tier model.
- cyanydeez 5d agoyou mean like a Shepards Tone (https://en.wikipedia.org/wiki/Shepard_tone https://en.wikipedia.org/wiki/Shepard_tone); i wouldn't doubt they slowly tweak quants to try to eke out. there's also probably load balancers that downgrade models during high use.