4 ms·
> but then why are we comparing such distinct and unrelated tasks as ... Because a few years ago the LLMs could only do trivial tasks that a child could do, an
by jsnell 1y ago
> but then why are we comparing such distinct and unrelated tasks as ...
Because a few years ago the LLMs could only do trivial tasks that a child could do, and now they're able to do complex research and software development tasks.
If you just have the trivial tasks, the benchmark is saturated within a year. If you just have the very complex tasks, the benchmark is has no sensitivity at all for years (just everything scoring a 0) and then abruptly becomes useful for a brief moment.
This seems pretty obvious, and I can't figure out what your actual concern is. You're just implying it is a flawed design without pointing out anything concrete.
- danlitt 1y agoThe key word is "unrelated"! Being able to count the number of words in a paragraph and being able to train an image classifier are so different as to be unrelated for all practical purposes. The assumption underlying this kind of a "benchmark" is that all tasks have a certain attribute called complexity which is a numerical value we can use to discriminate tasks, presumably so that if you can complete tasks up to a certain "complexity" then you can complete all other tasks of lower complexity. No such attribute exists! I am sure there are "4 hour" tasks an LLM can do and "5 second" tasks that no LLM can do. The underlying frustration here is that there is so much latitude possible in choosing which tasks to test, which ones to present, and how to quantify "success" that the metrics given are completely meaningless, and do not help anyone to make a prediction. I would bet my entire life savings that by the time the hype bubble bursts, we will still have 10 brainless articles per day coming out saying AGI is round the corner.
- Voultapher 1y agoWell put, the metric is cherry picked to further the narrative.