3 ms·
Have they improved? Is there evidence of that? Got a task you were doing in 2024 and 2026 and the results of each?
by pocksuppet 1mo ago
Have they improved? Is there evidence of that? Got a task you were doing in 2024 and 2026 and the results of each?
- threatripper 1mo agoThere is plenty of evidence that they have improved in all benchmarks and also in my private experience. But have they improved in the things they still fail at? No, they still fail at them. You need only one example of failure to prove that it still fails. They still fail a lot on many real world tasks. So, depending on what you ask, they may have not improved even a tiny bit.
- Denkel 29d agoI am a very light user, so this is my feeling from reading about other people's experience; I wouldn't say that they plateau'd but up to 4.5/4.8 the gains in the models felt exponential while since then they feel more linear and the big improvements are coming less from the models and more from everything around it (harnesses, agentic development, skills...). So, while I don't feel like there has not been improvement, it really feels like there is a limit that will be reached sooner than later (and for sure before any AGI).
- anthonyrstevens 1mo agoOh my goodness. Is this really a good-faith question?