4 ms·
These tests are pointless when often models get nerfed few days after release. Astra today is way dumber than just few days ago.
by varispeed 23d ago
These tests are pointless when often models get nerfed few days after release. Astra today is way dumber than just few days ago.
- cbg0 23d agoNot Astra but Sol hasn't been nerfed: https://marginlab.ai/trackers/codex/ https://marginlab.ai/trackers/codex/
- d_tr 23d agoHow and why do they get nerfed? To save money?
- varispeed 23d agoYes. They save on compute and customer has to use more tokens to achieve their goal which means more profit.
- cliche 23d agoProfit? I thought these companies were making a massive loss
- johnfn 23d agoModels do not get nerfed. There has never been evidence of this. This would be trivial to prove if it were true, and such a proof would be a huge story and scandal to a news market hungry for a shred of a signal on AI's downfall. This is the "your iPhone is listening to you and serving ads based on what you say" of the 2020s.
- esafak 23d agoYes, they can, through quantization. Many providers of open source models openly serve quantized versions; check openrouter.
- 4chandaily 23d ago> This is the "your iPhone is listening to you and serving ads based on what you say" of the 2020s Of course, it turned out that this wasn't actually completely BS. We just were accusing the wrong vendor. Not disagreeing with you on models.
- johnfn 23d agoI'd be happy to read a source as I am fairly confident that audio transcription -> facebook (or other) ads has never been true.
- 4chandaily 23d agoI was referring to LG specifically. Here is a link from https://news.ycombinator.com/item?id=49592375 https://news.ycombinator.com/item?id=49592375 three days ago.
- johnfn 23d agoWow totally missed this, thanks for sharing
- ricardobeat 23d agoOne theory is that they start serving at full precision, then quantize to save on costs as adoption grows. It's kind of a conspiracy theory atm, but I have definitely felt it - I had a large project done on Opus 4.8 release day, a week later it was struggling to complete partial tasks in the same area.
- kzrdude 23d agoIs that something we have credible evidence for? Do they serve a better model when artificialanalysis (the benchmark site) is making the requests, and so on?
- varispeed 23d agoIt is my own anectodal. Few days something I worked on usually got one shotted or got quality result. Today it is very much going nowhere and is stuck in reasoning loops. Probably someone should build nerf tracker, because this is quite common that models get substantially worse once PR hype wears off and they quantise them more or simply route requests to older models with system prompt changed to say it is Astra and not Sol etc.
- EPWN3D 23d agoSo it's your own anecdotal experience, but someone should build a service to track it?
- singingtoday 23d agoThere is a service to track anthropic models. Not sure how accurate it is, but my vibe says somewhat.