3 ms·
Even if it is a joke, having a consistent methodology is useful. I did it for about a year with my own private benchmark of reasoning type questions that I alwa
by fzzzy 1y ago
Even if it is a joke, having a consistent methodology is useful. I did it for about a year with my own private benchmark of reasoning type questions that I always applied to each new open model that came out. Run it once and you get a random sample of performance. Got unlucky, or got lucky? So what. That's the experimental protocol. Running things a bunch of times and cherry picking the best ones adds human bias, and complicates the steps.
- simonw 1y agoIt wasn't until I put these slides together that I realized quite how well my joke benchmark correlates with actual model performance - the "better" models genuinely do appear to draw better pelicans and I don't really understand why!
- MichaelZuo 1y agoI imagine the straightforward reason is that the “better” models are in fact significantly smarter in some tangible way, somehow.
- pama 1y agoHow did the pelicans of point releases of V3 and of R1 (R1-0528) do compared to the original versions of the models?
- more-nitor 1y agoI just don't get the fuss from the pro-LLM people who don't want anyone to shame their LLMs... people expect LLMs to say "correct" stuff on the first attempt, not 10000 attempts. Yet, these people are perfectly OK with cherry-picked success stories on youtube + advertisements, while being extremely vehement about this simple experiment... ...well maybe these people rode the LLM hype-train too early, and are desperate to defend LLMs lest their investment go poof? obligatory hype-graph classic: https://upload.wikimedia.org/wikipedia/commons/thumb/9/94/Gartner_Hype_Cycle.svg/559px-Gartner_Hype_Cycle.svg.png https://upload.wikimedia.org/wikipedia/commons/thumb/9/94/Ga...
- tuananh 1y agountil they start targeting this benchmark
- simonw 1y agoRight, that was the closing joke for the talk.
- jonstewart 1y agoIt is funny to think that a hundred years in the future there may be some vestigial area of the models’ networks that’s still tuned to drawing pelicans on bicycles.
- johnrob 1y agoWell, the most likely single random sample would be a “representative” one :)
- famouswaffles 1y agoLLMs also have a 'g factor' https://www.sciencedirect.com/science/article/pii/S0160289624000527 https://www.sciencedirect.com/science/article/pii/S016028962...
- Breza 1y agoAnother advantage is you can easily include deprecated models in your comparisons. I maintain our internal LLM rankings at work. Since the prompts have remained the same, I can do things like compare the latest Gemini Pro to the original Bard.