4 ms·
I made some LLM-powered text-adventure games: https://cosmictrip.space/gameannouncement https://cosmictrip.space/gameannouncement And I'm working on a webapp t
by cryptoz 3y ago
I made some LLM-powered text-adventure games: https://cosmictrip.space/gameannouncement https://cosmictrip.space/gameannouncement
And I'm working on a webapp that is a kanban board where LLM and human collaborate to build features in code. I just got a cool thing working there: like everyone, having LLM generate new code is easy but modifying code is hard. So my attempt at working on modifying code with LLM is starting with HTML and having GPT-4 write beautfulsoup code that then makes the desired modification to the HTML file. Will do with js, python via ast, etc. No link for this one yet :) still in development.
- s-macke 3y agoI didn't make text-adventures with LLMs. I try to solve them [0]. So, far, none of the 7 tested models were able to win even one of the easiest text adventures. I tried many prompting techniques. But only GPT-4 was able to play through the first half of the game. [0] https://github.com/s-macke/AdventureAI https://github.com/s-macke/AdventureAI
- ianbicking 3y agoFun, I tried to do this back with GPT-3: https://llm.ianbicking.org/interactive-fiction/ https://llm.ianbicking.org/interactive-fiction/ But Zork wouldn't be a very accurate measure of skill because GPT definitely knows Zork. Unfortunately the emulator (https://github.com/DLehenbauer/jszm https://github.com/DLehenbauer/jszm) doesn't work with most games newer than Zork. I haven't revisited the code with newer GPT models either.
- s-macke 3y agoGPT-3 doesn't even manage the first few steps of the tested text adventure. And GPT-4 is not good at playing these adventures either. However, my code run a newer version of the Z-machine. So Zork and many other text adventures will work. I have not tried many other games though.
- ianbicking 3y agoI was surprised how high your costs were. I assume you are putting the entire transcript into each prompt, but even then that seems high. Is GPT's planning also taking up a lot of room? I did find giving GPT some hints about the known commands helped a lot, and I put in some detection of error messages and kept a running log of commands that wouldn't work. Getting it to navigate the parser is kind of half of the skill of playing one of these games. It would be interesting to have it play some, then step back and have it reflect and enumerate things about how the play itself works.
- s-macke 3y agoThe costs have dropped significantly months after I created the cost image. Now I use GPT-4 Turbo. This GPT-4 model understand how text adventures work and there is no need to give him known commands. Of course you try even more sophisticated techniques than mine. I tried the ReAct pattern and virtual discussions. So far, he always stumbles at the same place in a critical understanding of the text. And I tried exactly this critical step dozens of times. You will understand the issue yourself, once you play the game yourself. It just takes 20 minutes and is very easy: https://adamcadre.ac/if/905.html https://adamcadre.ac/if/905.html
- ianbicking 3y agoYou mean at the very end of the game? The game seems like it's only designed to trick you into that very ending :) Are you hoping it will figure out the game based on the context clues? I'm not sure I can find them myself... A long time ago I did some exercises in "classical planning algorithms", which all feel very like the early part of this game. I.e., how do you get ready to leave if you have to shower, and can't do that with clothes on, etc. A similar planning example involved changing a tire (opening the trunk, removing lug nuts, etc). It was surprisingly difficult to make an algorithm that could figure it out! You could search the state space given the transitions, but it exploded with what was effectively lots of dead ends; obvious to me as a human, but not to the algorithm. Which is to say that this is a harder problem than it might seem.