4 ms·
If benchmark means anything in evaluating how model capability progress, the evidence is that all the existing benchmark have been pretty much solved, except Fr
by nerder92 2y ago
If benchmark means anything in evaluating how model capability progress, the evidence is that all the existing benchmark have been pretty much solved, except FrontierMath (https://epoch.ai/frontiermath https://epoch.ai/frontiermath)
- shihab 2y agoI recently saw Sam Altman bragging about OpenAI's performance on Codeforces (leetcode-like website), which I consider just about the worst benchmark possible. 1. All problems are small- the prompt and solution (<100 LOC, often <60LOC) 2. Solving those problems is more about recollecting patterns and less about good new insights. Now, top level human competitors do need original thinking, but that's only because our memory is too small to store all previously seen patterns. 3. Unusually good dataset- you have tens of thousands of problems, each with thousands of submissions, along with clear signals to train on (right/wrong, time taken etc), a very rich discussion sections etc. I think becoming 100th best Codeforces programmer is still an incredible achievement for a LLM. But for Sam Altman to specifically note the performance on this- I consider that a sign of weakness, not strength.
- badgersnake 2y agoAltman spouts even more bullshit than his models, if that’s even possible.
- layer8 2y agoGiven that I’ve seen excellent mathematicians produce poor-quality code in real-world software projects, I’m not sure how relevant these benchmarks are.
- mrguyorama 2y agoBenchmarks cannot tell you whether the tech will continue supernaturally, linearly, or plateau entirely. In the 90s, companies showed graphs of CPU frequency and projected we would be hitting 8ghz pretty soon. Futurists predicted we would get CPUs running at tens of ghz. We only just now have 5ghz CPUs despite running at 4ghz back in the mid 2000s. We fundamentally missed an important detail that wasn't consider at all in those projections. We know less about the theory of how LLMs and neural networks grow with effort than we did about how transistors operate over different speeds. You utterly cannot extrapolate from those kinds of graphs.