4 ms·
> In conclusion most of the Large Language Models can play and win text adventures Doesn't the table above show that NONE of the LLMs won the game, not even GP
by singularity2001 2y ago
> In conclusion most of the Large Language Models can play and win text adventures
Doesn't the table above show that NONE of the LLMs won the game, not even GPT-4?
- lolinder 2y agoThis game specifically is meant to not be won in the first pass. Any model that completes this game once would likely win simpler games that are only meant to take one pass. What the table shows is that the top three models were able to finish the game and get the twist ending but were unable to carry forward information from the failed run into a second run. That said, I'd actually be curious to see what happens if Claude were coerced into continuing the game instead of bailing. With its large context window I think it'd have the best chance at actually figuring it out, so it's unfortunate that the author accepted its resignation.
- s-macke 2y agoOf course I have forced Claude to continue. But it seems, that Claude had everything forgotten from the bad ending. SITUATION: I have just woken up in a messy bedroom to a ringing phone. It seems I overslept after a difficult night, and the phone ringing urgently suggests that something important and time-sensitive is happening. THOUGHT: The incessant ringing of the phone is very ominous. Given the state of the room and my oversleeping, I have a bad feeling that I am extremely late for something critical, like an important work obligation. I should answer the phone immediately to find out the situation and how much trouble I am in. But I'll also grab my wallet and keys from the end table first in case I need to rush out the door right after the call. COMMAND: take wallet and keys
- s-macke 2y agoWell, technically yes. Three models so far have managed to finish the first bad ending. But then it seems, all models fail at this specific point. I had hundred of reruns of exactly that bad ending and tried different prompting techniques. So far, no luck. None of them ever looked under the bed. Anyhow, the results show, that it can handle such an easy adventure. And not all of the text adventures have such a difficult twist.