3 ms·
Once again, I must ask everyone not to place too much emphasis on this benchmark. Another post I see right now on the HN homepage is this: https://www.tbray.org
by zone411 2y ago
Once again, I must ask everyone not to place too much emphasis on this benchmark. Another post I see right now on the HN homepage is this: https://www.tbray.org/ongoing/When/202x/2024/04/18/Meta-AI-oh-my https://www.tbray.org/ongoing/When/202x/2024/04/18/Meta-AI-o.... As it points out, Llama 3 gave a plausible, smart-sounding answer and people would rate it highly on the LMSYS leaderboard, yet it might be totally incorrect. It's best to think of the LMSYS ranking as something akin to the Turing Test, with all its flaws.
That said, all other benchmarks so far (including my NYT Connections benchmark) show that both Llama 3 models are exceptionally strong for their sizes.