3 ms·
Interesting that in the video, there is an admission that they have been targeting this benchmark. A comment that was quickly shut down by Sam. A bit puzzling
by roboboffin 2y ago
Interesting that in the video, there is an admission that they have been targeting this benchmark. A comment that was quickly shut down by Sam.
A bit puzzling to me. Why does it matter ?
- HarHarVeryFunny 2y agoIt matters to extent that they want to market this as general intelligence, not as a collection of narrow intelligences (math, competitive programming, ARC puzzles, etc). In reality it seems to be a bit of both - there is some general intelligence based on having been "trained on the internet", but it seems these super-human math/etc skills are very much from them having focused on training on those.
- roboboffin 2y agoHowever, the way it is progressing is that the SOTA is saturating the current benchmarks; then a new one is conceived as people understand the nature of what it means to be intelligent. It seems only natural to concentrate on one benchmark at a time. Francois Chollet mentioned that the test tries to avoid curve fitting (which he states is the main ability of LLMs). However, they specifically restricted the number of examples to do this. It is not beyond the realms of possibility that many examples could have been generated by hand though, and that the curve fitting has been achieved, rather than discrete programming. Anyway, it’s all supposition. It’s difficult to know how genuine the results is, without knowledge of how it was actually achieved.
- mukunda_johnson 2y agoI always smell foul play from Sam. I'd bet they are doing something silly to inflate the benchmark score. Not saying they are, but Sam is the type of guy to put a literal dumb human in the API loop and score "just as high as a human would."