36 ms·
IMO Lmsys benchamarks are essentially: How AI nerds (like us, so a particular subclass of the "normal" population) rate responses from various AIs - this ratin
by scrollop 2y ago
IMO Lmsys benchamarks are essentially:
How AI nerds (like us, so a particular subclass of the "normal" population) rate responses from various AIs - this rating is subjective, of course, and an LLM that produces answers that read well though is less accurate/"intelligent" would likely perform better than a model that propduces answers that read worse, though is more "intelligent".
AI nerds/IT nerds likely have higher prevalence of ASD/Asperger's etc (or higher up the spectrum), so likely not the best way to rank LLMs.
It's fun, though!