3 ms·
Chatbot Arena: Benchmarking LLMs in the Wild with Elo Ratings
- HoshinoAI 3y agoIt's glad to see the old technique is used for new models. You may also learn from 7.33 dota update which uses a new ranking algorithm called Glicko.
- zhisbug 3y agocould you provide any reference? Is it a variant of ELO?
- bradknowles 3y agoElo != ELO. One is a rating system named after the creator, Arpad Elo. See https://en.wikipedia.org/wiki/Elo_rating_system https://en.wikipedia.org/wiki/Elo_rating_system The other is a rock band that was formed in 1970. See https://en.wikipedia.org/wiki/Electric_Light_Orchestra https://en.wikipedia.org/wiki/Electric_Light_Orchestra
- HoshinoAI 3y agocheck matchmaking section: https://www.dota2.com/newfrontiers https://www.dota2.com/newfrontiers Valve listed some reason for making the change. https://en.wikipedia.org/wiki/Glicko_rating_system https://en.wikipedia.org/wiki/Glicko_rating_system
- weichiang 3y agoSurprised to learn StableLM is worse than plain LLaMA. link to their leaderboard: leaderboard.lmsys.org
- circuit10 3y agoI’ve heard that it’s really bad for it’s size
- lee101 3y ago[dead]
- freediver 3y agoThis is a very good idea.
- aaron695 3y ago[dead]