69 ms·
> I can also describe a new/novel game or puzzle to GPT4 and it can have a go at playing and solving it. It’s interesting that you brought up chess. It can do
by kweingar 3y ago
> I can also describe a new/novel game or puzzle to GPT4 and it can have a go at playing and solving it.
It’s interesting that you brought up chess. It can do chess reasonably because there is a huge amount of chess data on the web. In that sense, it is not too surprising to me. If someone several years ago had said “I scoured the entire internet for chess-related text and fed it into an AI model, and it can play at a low-amateur level” I would be impressed but I wouldn’t be hailing a new era of general intelligence.
An example that illustrates that huge amounts of specialized data is needed for it to do any particular task: I fed GPT-4 the rules of Duck Chess. Duck Chess is exactly like regular chess but after each move, the player who just moved takes the rubber duck and places it on any empty square (there is just one duck shared between the players, and you have to move the duck: you can’t leave it on the same square for consecutive moves). Pieces cannot move through or stop on the duck. The game eliminates the concepts of check and checkmate, and ends when a player captures the opposing king.
I have given a description of duck chess to many humans (yes, I love duck chess!) who are usually much worse than ChatGPT at regular chess. When these humans play duck chess for the first time, they intuit some basic principles: use the duck to block natural developing moves for your opponent in the opening; you can often capture a defended piece without consequence by placing the duck between your capturing piece and the defender; if you want the duck to not be on a certain square for your next move, then put it on that square after your own move, since your opponent is obligated to move it; and so on.
GPT-4 meanwhile utterly fails to play the game. More often than not, it will try something illegal: putting the duck on an occupied square, passing a piece through the duck as though it weren’t there, or attempting to capture the duck after being told that’s not possible. When it does play legal moves, the duck placement is nonsensical. When asked why it placed the duck where it did, it betrays a lack of basic understanding of the rules. Its explanations tend to forget that its opponent gets to move the duck themself after their move.
This is where the “but humans make mistakes too!” arguments break down. No human who can play regular chess at the level of GPT-4 would continually struggle to make legal moves in duck chess. 99.999% of them would make better moves than GPT-4.
To me, this supports the idea that GPT-4 is great at finding and exploiting patterns that it has seen millions of times in the training set. When you veer off the training data (and your problem isn’t a trivial interpolation of related concepts that are in the training data) it seems to fall apart completely.
As one more example, GPT-4 contains some very basic facts about the game Arimaa, an abstract strategy game like chess. It can recite the rules perfectly. But I can’t play Arimaa with GPT-4 because it fails on the very first step: choosing how to arrange your pieces. I once exhausted all 25 of my messages trying to get it to make a legal configuration of its pieces to start the game, to no avail.