3 ms·
That's a nice experiment! I think codenames could definietly be an evaluation method for LLMs.
by suveen_ellawela 2y ago
That's a nice experiment! I think codenames could definietly be an evaluation method for LLMs.
- pieix 2y agoElo on different card games/board games would be a great eval metric now that the systems are general enough to play Codenames, chess, poker…
- suveen_ellawela 2y agototally agree!
- __MatrixMan__ 2y agoIt would be fun to build one, perhaps mediated by an app, where you have to guess whether your spymaster is a human or an AI based on the quality of their choices.
- zeroonetwothree 2y agoThe average human is quite bad. It really works well when the spymaster is (a) experienced and (b) familiar with the other players.
- __MatrixMan__ 2y agoIt's the (b) case I'm interested in. Like the spymaster loses if they can't subtly indicate to their friends that they're the real deal. Otherwise the robots win.
- suveen_ellawela 2y agoi thought of adding a feature where you can get your own spy master. you can give it all your personal info and the clues would be customized. the botteleneck is the other human spymaster has to help with updating the game state cus I(guesser) can't look at the spy master view.