2 ms·
>> "LLMs have reached a plateau." you should look at benchmarks such as ARC which went from "needs 10 years, currently at 0%" to almost solved within the leas
by singularity2001 10mo ago
>> "LLMs have reached a plateau."
you should look at benchmarks such as ARC which went from "needs 10 years, currently at 0%" to almost solved within the least year. Also there is a revolution happening in math which the layman might be missing.
- bloppe 10mo agoI don't care about the benchmarks. I care about how helpful coding agents are for my work. And I can barely tell the difference between the models this year and the models last year. Everyone's raving about Opus but I bet about 50% of people would be able to identify it in a blind test against Sonnet.
- jononor 10mo agoFor ARC v1 it was found that it was much less resistant to brute force than intended/designed. This was improved in v2, which LLMs are currently doing less good at. Note also that ARC tasks are explicitly designed to be slightly-out-of-reach, things that are quite simple for humans, but current models are pretty bad at - designed to measure and enable progress. But yeah there are many interesting approaches, and ARC is interesting to follow both because it attempts to measure ability to adapt to new takas ("fluid intelligence"), and because we have not saturated it yet.