4 ms·
I'd love to see a benchmark that tests different LLMs for slop, not necessarily limited to code. That might be even more interesting than ARC-AGI.
by growdark 11mo ago
I'd love to see a benchmark that tests different LLMs for slop, not necessarily limited to code. That might be even more interesting than ARC-AGI.
- jampa 11mo agoNot a benchmark per se, but there is a "Not x, but y" Slop Leaderboard: https://www.reddit.com/r/LocalLLaMA/comments/1lv2t7n/not_x_but_y_slop_leaderboard/ https://www.reddit.com/r/LocalLLaMA/comments/1lv2t7n/not_x_b...
- Bolwin 11mo agoSee the writing benchmarks here https://eqbench.com/creative_writing_longform.html https://eqbench.com/creative_writing_longform.html
- Der_Einzige 11mo agoNote this is the same first author
- deleted 11mo ago[deleted]
- topaz0 11mo ago100% of LLM output is slop. Done.