4 ms·
> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better
by AmazingTurtle 10d ago
> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?
Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out.
Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.
- ramesh31 10d ago>Also I will likely save some money on subscriptions. Unlikely. The $200 Claude subscription allows for billions of tokens/month, and that kind of hardware will take years to amortize.
- torben-friis 10d agoThere could be gym logic at play. Hundreds signed up, 20 people actually exercising. Though it's probably more likely in the lower tiers.
- rybosworld 10d agoI wouldn't be so sure. The generosity of the subscription plans has declined GREATLY over the past 6 months or so. They are likely trending towards api pricing parity. In which case, having your own hardware makes sense if you can utilize it well.
- zeroonetwothree 10d agoLast time I estimated it was like 30 years to pay back. I doubt the hardware will even last that long.
- fragmede 10d agoLast time I estimated, it would only take 3 months to pay back because the 1TB Mac Mini running Qwen RSIingly developed ASI and made infinity dollars off of crypto and I got put in jail by the SEC. Where'd you get 30 years from? Show your work.
- wilj 10d agoI would like to subscribe to your newsletter.
- mike_d 10d agoI have 2 x ChatGPT Pro 20x, Claude Max 20x, and Kimi Vivace. It's about ~12 months payback for two units and the cable. The problem is they can't fit any frontier level open models.
- AtHeartEngineer 10d agoflash next is good, I've been running it for like 2 weeks now and it's pretty solid, hope you like it and it meets your needs. I still lean on Claude and codex a fair bit for harder stuff, but I'm rapidly moving towards 2x $20 plans instead of 2x $200 plans
- boardwaalk 10d agoI don’t know what people do with the open models but having tried a lot of them I just can’t make it make sense. they’re too dumb and it effectively makes them useless (to me). it’s probably worth being honest about the low ceiling here.
- cyanydeez 10d agoQwen3.8-Flash-Next seems pretty much auto pilot when I get it the right context. Perhaps reverse the question: Are your build/construct requirements just really counter-productive to how LLMs need to understand things? I've found constructing the code, writing the tests, adding the docs; then running through them gets most of the way there. I've also found that making a simple obvious edit is a useless endevour when the LLM is primed for the long context tasks. So, again, the question is reversed: are you over reliant on the LLM to do even stupid simple likes like editting a css variable?
- redanddead 10d agoServing compute is their main value prop Yet… even Altman called out Anthropic for serving dumbed down models. Shits weird man