7 ms·
OpenAI's o1 Playing Codenames
- JaggerFoo 2y agoI did this with Claude over the holidays. Putting Claude in the role as a guesser and comparing the guess to another experience human player. It turns out they both matched each other.
- suveen_ellawela 2y agoThat's a nice experiment! I think codenames could definietly be an evaluation method for LLMs.
- pieix 2y agoElo on different card games/board games would be a great eval metric now that the systems are general enough to play Codenames, chess, poker…
- suveen_ellawela 2y agototally agree!
- __MatrixMan__ 2y agoIt would be fun to build one, perhaps mediated by an app, where you have to guess whether your spymaster is a human or an AI based on the quality of their choices.
- zeroonetwothree 2y agoThe average human is quite bad. It really works well when the spymaster is (a) experienced and (b) familiar with the other players.
- __MatrixMan__ 2y agoIt's the (b) case I'm interested in. Like the spymaster loses if they can't subtly indicate to their friends that they're the real deal. Otherwise the robots win.
- suveen_ellawela 2y agoi thought of adding a feature where you can get your own spy master. you can give it all your personal info and the clues would be customized. the botteleneck is the other human spymaster has to help with updating the game state cus I(guesser) can't look at the spy master view.
- joaomacp 2y agoI tried whatever the multi-modal paid ChatGPT model is on the Codenames Pictures version, and it didn't fare that well. Since they will probably scrape this comment and add it to next model's training data, I look forward to it getting good!
- suveen_ellawela 2y agohaha!
- kennyloginz 2y agoCould this just be a case of Reddit being included in the training data? “ I read through codenames official rules to see if using "007" as a clue was allowed, and it turns out it is! To my surprise, I even came across a Reddit post where people were discussing and justifying why this clue fits perfectly within the rules.”
- JohnMakin 2y agoYea, initially I thought this post was satire because of this.
- suveen_ellawela 2y agothat is a really interesting point. if it is true, this shows direct usage of a single training data point ( cus there are no other resources talking about this fact)
- tsroe 2y agoFun quirk about this game: If there aren't too many cards left and your teammate knows their powers of two, you have a winning strategy. You simply lay a mental bitmap over all remaining cards, setting 1 for cards that belong to your team and 0 for all others. You can then just say the number that is represented by this bitmap, e.g. "five" for 0101, and your teammate can decode it in their head. All numbers are, after all, single words. This means, if you are very good at mental maths or you allow for a calculator, you could also win every game in the first round. For me personally however, it only becomes feasible with around 10 cards remaining.
- RedNifre 2y agoThat's against the rules.
- deleted 2y ago[deleted]
- Klaster_1 2y agoGuys I was playing with declared a similar move against the rules, so it was back to the old latent space search.
- Smaug123 2y agoIt is explicitly against the rules (https://czechgames.com/files/rules/codenames-rules-en.pdf https://czechgames.com/files/rules/codenames-rules-en.pdf), so they were correct. "Your clue must be about the meaning of the words. You can't use your clue to talk about the letters in a word or its position on the table."
- andrepd 2y agoThis is explicitly against the rules.
- tweakimp 2y agoWhat if the game showed a different order of cards to every player?
- thrance 2y agoI mean, it's playing against itself, not really a fair comparison to humans in my mind. The fun and hard part of this game is to get into your teammates brains and decipher what they possibly meant with what they played.
- suveen_ellawela 2y agoyea, didn't mean to take the fun out of the original game. the idea for this came when we asked chatgpt how to connect the words 'carrot' and 'ray'. maybe you can give a try too!
- thrance 2y agoI still enjoyed reading your post, it's fun and interesting! Maybe one could try having two different models play together, to see if they are genuinely good at the game or simply able to infer their own reasoning, if that makes sense. I'm kinda bad at word games like codenames, even in my native language (french). With carrot and ray, I'd try something like "striation"? But it's really convoluted.
- xnickb 2y agoSomehow I expected AI to give clues that combine 4-5-6 words at a time. It's not at all impressive to me. And I'm not a serious player at all
- pama 2y agoI was wondering about the same. It is possible that the instructions didn’t try to make the gameplay as aggressive as possible. A good model could optimize the separator to make it easy to guess the most words possible. By having access to its own state, it should be possible to reach 5–6 words in most cases. There is an argument for keeping words around that would increase the difficulty of the opponents guessing large/clean separations, so it is possible that optimal play includes simple pairs on occasion. Very interesting application nonetheless.
- vitus 2y ago> It is possible that the instructions didn’t try to make the gameplay as aggressive as possible. In case you're wondering, the prompts are available here: https://github.com/SuveenE/codenames-ai/blob/main/utils/prompts.ts https://github.com/SuveenE/codenames-ai/blob/main/utils/prom...
- pama 2y agoThanks!
- vitus 2y agoI am similarly less-than-impressed. If you click through to the website, you can watch the replay of one of the games mentioned in the article (the one with the clue "invader"). In that instance, the clues all matched 2-3 words, and the winning team got lucky twice (they guessed an unclued word using an unintended correlation, and their opponent guessed a different one of their unclued words.) You also see a number of instances where the agents continue guessing words for a clue even though they've already gotten enough matches. For instance, in round 2, for the clue "Japan (2)", the blue team guesses sumo and cherry, then goes for a rather tenuous followup guess for round 1's 007 with "ring" (despite having gotten the two clued matches in the first round). A sillier example is in the final round, where the Red Team guesses 3 clues (thereby identifying all nine of their target words), then going ahead and guessing another word. (For what it's worth, I think "shark" would have been a better guess for another 007 tie-in seeing as there are multiple Bond movies with sharks, but it's also not a match, and again, I wouldn't have gone for a third guess here when there were only two clued words.)
- croes 2y agoIs that really surprising? It’s basically the same brain playing with itself. Seems quite natural to link the code names to the same words. Let different LLMs play.
- deredede 2y agoThis is the take I thought I'd have, but in the last example, the guesser model reaches the correct conclusion using a different reasoning than the clue giver model. The clue giver justifies the link of Paper and Log as "written records", and between Paper and Line as "lines of text". But the guesser model connects Paper and Log because "paper is made from logs" (reaching the conclusion through a different meaning of Log), and connects Paper and Line because "'lined paper' is a common type of paper". Similarly, in the first example, the clue giver connects Monster and Lion because lions are "often depicted as a mythical beast or monster in legends" (a tenuous connection if you ask me), whereas the guesser model thought about King because of King Kong (which I also prefer to Lion).
- unlikelymordant 2y agogenerally there is a "temperature" parameter that can be used to add some randomness or variety to the LLMs outputs by changing the likelihood of the next word being selected. This means you could just keep regenerating the same response and get different answers each time. each time it will give different plausible responses, and this is all from the same model. This doesn't mean it believes any of them, it just keeps hallucinating likely text, some of which will fit better than others. It is still very much the same brain (or set of trained parameters) playing with itself.
- suveen_ellawela 2y agoI wanted to play around with the temperature, but unfortunately o1 only supports '1' as the value.
- wizzwizz4 2y ago> But the guesser model connects Paper and Log because "paper is made from logs" (reaching the conclusion through a different meaning of Log) No, it doesn't. It reaches the conclusion because of vector similarity (simplified explanation): these explanations are post-hoc.
- fercircularbuf 2y agoI've intuitively felt that this general class of task is what these LLMs are absolutely best at. I'm not an expert on these things, but isn't this thanks to word embeddings and how words are mapped into high dimensional vector space within the model? I would imagine that because every word is mapped this way, finding a word that exists in the same area as mail, lawyer, log, and line in some vector space would be trivial for the model to do, right?
- infinitifall 2y agoMore than just words. I've found LLMs immensely helpful for searching through the latent space or essence of quotes/books/movies/memes. I can ask things like "whats that book/movie set in X where Y happens" or "whats that quote by a P which goes something like Q" in my own paraphrased way and with a little prodding, expect the answer. You'd have no luck with traditional search engines unless someone has previously asked a similar question.
- captn3m0 2y agoI've been trying to do this with just word2vec, instead of throwing an LLM, since you just need to find a word with the appropriate distances optimized. https://github.com/captn3m0/ideas?tab=readme-ov-file#codenames-ai https://github.com/captn3m0/ideas?tab=readme-ov-file#codenam...
- dartos 2y agoI love this. Imagine the energy savings if more people didn’t just automatically reach for LLMs for their pet projects.
- zeroonetwothree 2y agoI tried this many years ago (before LLMs) with hundreds of real human games and it was never that good.
- qqqult 2y agoI did that last summer, I compared the performance of different english word embedding models, as far as I remember the best ones were GloVe and a few knowledge graph word embeddings. None of them were better than a human at giving hints for 3+ words though
- tweakimp 2y agoIt would be really interesting to see an LLM watch other players and learn how they think to find the best clues THEY need to hear to find the right words.
- suveen_ellawela 2y agodefinietly an interesting approach. I started writing down gameplays when i play with friends. then eventually stopped to enjoy the moment.
- progrus 2y agoGPT-3 was superhuman at this too
- suveen_ellawela 2y agoyep, agree. One big part of the experiment was to see how well it does the reasoning by asking it to output the reasoning.
- sylware 2y agoIf it can port c++ to C99+ and write correct 64bit risc-v assembly...
- jprete 2y agoCodenames is absolutely dead-center of what I expect Large Language Models to be good at. The fundamental skills of the game are: having an excellent embedding for word semantics and connotations; modeling other people's embeddings; a little bit of game strategy related to its competitive nature.
- badgersnake 2y agoOr just play with your friends?
- suveen_ellawela 2y agomy friends were bad clue givers. i just had to switch to ai.
- badgersnake 2y agoThat’s part of the fun, though.
- zeroonetwothree 2y agoI don’t find this “super good”. It’s mostly giving 2 clues which is the most basic level of competence. The paper 4 clue is reasonable but a bit lucky (eg Jack is also a good guess). I also don’t see it actually using tactics properly, which I would consider part of being “super good”. The game isn’t just about picking a good clue each round! Now obviously it’s still pretty decent at finding the clues. Probably better than a random human who hasn’t played much. Just I find the post’s level of hype overstated. It feels like the author isn’t very experienced with Codenames. It would be interesting to compare AI:human vs human:human games to see which does better. It seems like AI:AI will overstate its success.
- groggo 2y agoCan you elaborate on some of the more advanced tactics? When I play, it's mostly about getting a good 2 clue each time. Then if you can opportunistically get a 3 or 4, that's awesome. Some tactics come in for choosing the right pairs of 2's so you don't end up mismatched, or leaving clues that might be ambiguous with your opponent's... But that's mostly it. It'll be fun for multiplayer! Just like how in other online games you can add in a AI to play as one of the players.
- mtmickush 2y agoOther advanced tactics involve giving a broad clue that matches 3-4 of your own and just one other (either your opponents or a civilian). Your team can pick up all the matches across several turns and the one off doesn't hurt as much as the plus four helps
- hunter2_ 2y agoThe S-tier tactic: When that high-number clue is cut short by a turn-ending mistake, the guessers tell their clue giver to inflate the number given during the totally unrelated next clue by however many remained from the truncated turn for which they don't need additional information to locate (and therefore it would be wasteful for a future clue to re-group those) so the stated number of that next clue must allow for its own cards plus the prior cards. Example: The clue is "places 4" and the guessers choose 1 correctly and then 1 wrong answer, but they had achieved consensus about 2 others (and are confused about only the remaining 1). So the turns ends but they inform the clue giver to inflate by 2 next turn. That clue giver (after the other team goes) will then say the clue is "people 5" and the guessers will know that they shall select 2 places and 3 people. This can cascade beyond just a pair of turns.
- raphael123 2y ago[dead]
- raphael1234 2y ago[dead]
- jsemrau 2y agoI have been doing some experiments with Agents, Reinforcement Learnings playing a 4x4 Tic Tac Toe game.[1]. Given my analysis of the "thought" process we are still really far from true understanding of such games. While in my game as well as OP"s, the rules are pre-trained and the models are good enough to reach a conclusion (which in itself is already impressive), it is still a long way. [1] https://jdsemrau.substack.com/p/nemotron-vs-qwen-game-theory-and https://jdsemrau.substack.com/p/nemotron-vs-qwen-game-theory...
- lolinder 2y agoA small weakness in this test is that one of the keys to strategic Codenames play is understanding your partner. You're not just trying to connect the words, you're trying to connect them in a way that will be obvious to your partner. As a computing analogy: you're trying to serialize a few cards in a way that will be deserializable by the other player. This test pairs o1 with itself, which means the serializer is the deserializer. So while it's impressive that it can link 4 words, most humans could also easily link 4 with as much stretching! We just don't tend to because we can't guarantee that the other human will make the same connections we did.
- ModernMech 2y agolol I played this game with my family and they said my wife and I were cheating because I kept using inside jokes that made no sense to them but she would get immediately.
- dgritsko 2y agoThat's a big part of what makes this game enjoyable - a clue that is very obvious to one person might not even cross the mind of someone else. To anyone reading this who hasn't played, it's definitely worth giving it a try.
- slyn 2y agoAgreed, big fan of codenames in general but it plays its best when you’re playing against / alongside people that you’ve known for a while. The metagaming aspect of structuring clues to who your partner is really takes it to the next level.
- lupire 2y agoSame for Taboo for me. It's why we married.
- jncfhnb 2y agoEhhh I don’t think that’s accurate. The problem is not linking 4 words. It’s linking 4 words without accidentally triggering other, semantically adjacent words. This task could probably be solved nearly just as well with old school word 2 vec embeddings
- jerkstate 2y agoYou can pretty reliably get 2-clues and sometimes good 3-clues just using word2vec embedding similarity
- suveen_ellawela 2y agoagree. i think getting to have a look at o1's reasoning was pretty fun.
- simonw 2y agoI've been trying out various "reasoning" models (o1, R1, Gemini Thinking etc) against the NYT Connections word puzzle - it's a really interesting test of them. So far o1 Pro has been the most consistently successful: https://www.nytimes.com/games/connections https://www.nytimes.com/games/connections
- topaz0 2y agoWonder if they use llms to write those puzzles
- macromaniac 2y agoI made one where you play with the AI a few years back instead of AI v AI but never posted it anywhere if anyone wants to try, just updated it to gpt-4o-mini https://wordswithrobots.isotropic.us/ https://wordswithrobots.isotropic.us/
- blakeburch 2y agoLove the idea! Just wish you could clarify a number like you do in codenames. Otherwise, it just keeps going until all of its options are wrong.
- macromaniac 2y agoTrue, because then it feels more intentional (+ the extra strategy). It was definitely a bit thrown together- atm I only ever use it when I need a bit of practice before playing codenames.
- suveen_ellawela 2y agocool stuff!
- Amekedl 2y ago“o1 is more knowledgeable than the average human” “the toyota yaris can move faster than the average human” even opt-125m from years ago can pull more facts than the average human.
- bongodongobob 2y agoI played it with 3.5 and it was great. This isn't something o1 just picked up on.
- suveen_ellawela 2y agoyep, agree. One big part of the experiment was to see how well it does the reasoning by asking it to output the reasoning.
- lsy 2y agoSome of these clues wouldn't be very good for a human playing. "007" for example isn't a very good clue for "laser", not only because something happening to be in one of several films about a character doesn't rise to the typical level of salience, but also because other words on-board like "shark" and "astronaut" even moreso meet the criterion of featuring prominently in James Bond movies, and "astronaut" appears to be a game-ending choice.
- yantrams 2y agoI cracked myself up with a ridiculous train of thought for fun while playing Codenames once. It went a little something like this Star => Twinkle => Twinkle Khanna => Married to Akshay Kumar => Canadian Citizen => Maple Syrup ( Leaf ? )
- suveen_ellawela 2y agohaha, i've been in similar situations, but this one's something else.
- lispforlife 2y ago[dead]
- lynguist 2y agoI kinda have the same very subjective feeling where o1 is the first AI that is clearly superior to me.
- some_random 2y agoI don't find this remotely compelling, I can easily come up with clues that make sense to me to connect a ton of words the difficulty is coming up with clues that others will look at the same way. The last example is exactly what I mean, "paper" makes sense for those 4 only when you explain it. If "Line" counts then why not "Gum" (which is typically wrapped in paper) or if "Lawyer" is valid then why not "King" (who's decrees are written on what?).
- jinyang0220 2y agoDude I looooooooooved that game. How long did u spend building it?
- suveen_ellawela 2y ago2 weeks!