3 ms·
This paper only scratches the surface and feels incomplete, as it references only GPT-4 and mentions appendices that are not included. The examples are two year
by s-macke 1y ago
This paper only scratches the surface and feels incomplete, as it references only GPT-4 and mentions appendices that are not included. The examples are two years old.
For a more in-depth analysis of chatbots playing text adventures, take a look at my project. I haven’t updated it in a while due to time constraints.
[0] https://github.com/s-macke/AdventureAI https://github.com/s-macke/AdventureAI
- s-macke 1y agoThe answer to the paper's question is likely yes—especially if context is used effectively and memory and summaries are incorporated. In that case, chatbots can complete even more complex games, such as Pokémon role-playing games [0]. The challenge with benchmarking text adventures lies in their trial-and-error nature. It’s easy to get stuck for hundreds of moves on a minor detail before eventually giving up and trying a different approach. [0] https://www.twitch.tv/gpt_plays_pokemon https://www.twitch.tv/gpt_plays_pokemon
- duskwuff 1y ago> The challenge with benchmarking text adventures I'd argue that another major challenge is that many of the popular text adventures are described in copious detail by many online sources. A language model "playing" Adventure is likely to use the magic word XYZZY by rote, for example, rather than learning it from the game text.
- glimshe 1y agoI like your project because you try to compare the performance of different chatbots. At the same time, I certainly wouldn't say it's more complete than the paper - your landing page is somewhat superficial. Reading both is better than just reading either.