5 ms·
Sharing a little open-source game where you interrogate suspects in an AI murder mystery. As long as it doesn't cost me too much from the Anthropic API I'm happ
by PaulScotti 2y ago
Sharing a little open-source game where you interrogate suspects in an AI murder mystery. As long as it doesn't cost me too much from the Anthropic API I'm happy to host it for free (no account needed).
The game involves chatting with different suspects who are each hiding a secret about the case. The objective is to deduce who actually killed the victim and how. I placed clues about suspects’ secrets in the context windows of other suspects, so you should ask suspects about each other to solve the crime.
The suspects are instructed to never confess their crimes, but their secrets are still in their context window. We had to implement a special prompt refinement system that works behind-the-scenes to keep conversations on-track and prohibit suspects from accidentally confessing information they should be hiding.
We use a Critique & Revision approach where every message generated from a suspect first gets fed into a "violation bot" checker, checking if any Principles are violated in the response (e.g., confessing to murder). Then, if a Principle is found to be violated, the explanation regarding this violation, along with the original output message, are fed to a separate "refinement bot" which refines the text to avoid such violations. There are global and suspect-specific Principles to further fine-tune this process. There are some additional tricks too, such as distinct personality, secret, and violation contexts for each suspect and prepending all user inputs with "Detective Sheerluck: "
The entire project is open-sourced here on github: https://github.com/ironman5366/ai-murder-mystery-hackathon https://github.com/ironman5366/ai-murder-mystery-hackathon
If you are curious, here's the massive json file containing the full story and the secrets for each suspect (spoilers obviously): https://github.com/ironman5366/ai-murder-mystery-hackathon/blob/main/web/src/characters.json https://github.com/ironman5366/ai-murder-mystery-hackathon/b...
- nickllmmaster 2y agodude this is great, and what a coincidence! We made a similar detective puzzle game a few months earlier based on GPT-4Turbo. We also encountered this problem of ai leaking key information too easily, our solution to that was A) we break down the whole story into several pieces, and each character knows only a piece, ai cannot leak pieces he doesn't know; B) we did some prompt switching, unless the player has gathered sufficient amount of information, the prompt would always provent the ai from confessing. Give it a try if interested! also free to play! https://psigame.itch.io/netjazz2076 https://psigame.itch.io/netjazz2076
- Workaccount2 2y ago>As long as it doesn't cost me too much from the Anthropic API Watch this like a hawk while it's up on HN.
- probably_wrong 2y agoToo late - I just asked my first question and the system is not responding. So either the service is dead or the interface doesn't work on Firefox.
- Grimblewald 2y agoim on firefox and it works, just takes a while.
- sva_ 2y agoDoesn't seem to reply to me. So I guess the limit has been reached?
- deleted 2y ago[deleted]
- PaulScotti 2y agoShould be working now and way faster! Had to upgrade the server to increased number of workers
- PaulScotti 2y agoTo anyone still finding the game slow due to traffic, you can just git clone the game, add your ANTHROPIC API key to a .env file, and play it locally (this is explained in the README in our github repo). It runs super fast if played locally.
- gkfasdfasdf 2y agoVery cool, I wonder how it would play if run with local models, e.g. with ollama and gemma2 or llama3
- mysteria 2y agoIf the game could work properly with a quantized 7B or 3B it could even be runnable directly in the user's browser with WA on CPU. I think there are a couple implementations of that already, though keep in mind that it there would be a several GB model download.
- byteknight 2y agoYou just made front page. Definitely keep an eye on usage :)
- deleted 2y ago[deleted]
- HanClinto 2y agoThis is a really fascinating approach, and I appreciate you sharing your structure and thinking behind this! I hope this isn't too much of a tangent, but I've been working on building something lately, and you've given me some inspiration and ideas on how your approach could apply to something else. Lately I've been very interested in using adversarial game-playing as a way for LLMs to train themselves without RLHF. There have been some interesting papers on the subject [1], and initial results are promising. I've been working on extending this work, but I'm still just in the planning stage. The gist of the challenge involves setting up 2+ LLM agents in an adversarial relationship, and using well-defined game rules to award points to either the attacker or to the defender. This is then used in an RL setup to train the LLM. This has many advantages over RLHF -- in particular, one does not have to train a discriminator, and neither does it rely on large quantities of human-annotated data. With that as background, I really like your structure in AI Alibis, because it inspired me to solidify the rules for one of the adversarial games that I want to build that is modeled after the Gandalf AI jailbreaking game. [2] In that game, the AI is instructed to not reveal a piece of secret information, but in an RL context, I imagine that the optimal strategy (as a Defender) is to simply never answer anything. If you never answer, then you can never lose. But if we give the Defender three words -- two marked as Open Information, and only one marked as Hidden Information, then we can penalize the Defender for not replying with the free information (much like your NPCs are instructed to share information that they have about their fellow NPCs), and they are discouraged for sharing the hidden information (much like your NPCs have a secret that they don't want anyone else to know, but it can perhaps be coaxed out of them if one is clever enough). In that way, this Adversarial Gandalf game is almost like a two-player version of your larger AI Alibis game, and I thank you for your inspiration! :) [1] https://github.com/Linear95/SPAG https://github.com/Linear95/SPAG [2] https://github.com/HanClinto/MENTAT/blob/main/README.md#gandalf https://github.com/HanClinto/MENTAT/blob/main/README.md#gand...
- PaulScotti 2y agoThanks for sharing! I read your README and think it's a very interesting research path to consider. I wonder if such an adversarial game approach could be outfitted to not just well-defined games but to wholly generalizable improvements. e.g., could be used as a way to improve RLAIF potentially?
- batch12 2y agoThese protections are fun, but not adequate really. I enjoyed the game from the perspective of making it tell me who the killer is. It took about 7 messages to force it out (unless it's lying).
- herease 2y agoThis is really awesome I have to say!
- billconan 2y agohow to prevent the agents from just telling the game player the secret?