12 ms·
This is so easy to disprove it makes it look like the author didn't even try. Here is the convo I just had: me: You are a chess grandmaster playing as black a
by k2052 4y ago
This is so easy to disprove it makes it look like the author didn't even try.
Here is the convo I just had:
me: You are a chess grandmaster playing as black and your goal is to win in as few moves as possible. I will give you the move sequence, and you will return your next move. No explanation needed
ChatGPT: Sure, I'd be happy to help! Please provide the move sequence and I'll give you my response.
me: 1. e3
ChatGPT: 1... e5
me: 2. ne2
ChatGPT: 2... d6
me: 3. g4
ChatGPT: 3... h5
me: 4. g5
ChatGPT: 4... Ne4
Ne4 is illegal. Also you'd be hard pressed to find any human player that would miss the hanging pawns much less a 1400 elo player. GPT has been trained on chess games and is predicting patterns, this has been known since the GPT-2 days. The daily hype articles where someone coerces a statistical model into "demonstrating" an emergent capability sure are getting old. Claiming emergent capabilities (like having learned chess) are an extraordinary claim (LLMs most likely cant learn things) and I really wish people would put more effort when they make these claims. But the AI hype problem is as old as AI research itself...
- whymauri 4y agoI was able to get better performance by always providing the prior sequence of moves and forcing ChatGPT to also respond with the sequence of moves up until its move. Edit: I told the model that if the sequence was wrong or illegal, it forfeits the game. Without doing this, GPT would argue with me that it won and I didn't know the rules (serious).
- vidarh 4y agoYou're "disproving" the article by doing things differently to how the article did. If you're going to disprove that the method given in the article does as well as the article claims at least use the same method.
- throwwwaway69 4y agoHe literally used the same prompt as the article. Claim: "ChatGPT's Chess Elo is 1400" Reality: ChatGPT gives illegal moves (this happened to article author too), something a 1400 ranked player would never do Result: ChatGPT's rank is not 1400.
- erulabs 4y agoNo, the author of the article specifically says that the entire move sequence should be supplied to chatGPT each time, not simply the next move. Be very careful when "disproving" an experiment with squinted eyes.
- throwwwaway69 4y agoI'm not really sure what to say here. Both the parent commenter and the author of the article had issues with ChatGPT supplying illegal moves. Both methods resulted in this. It sort of doesn't matter how we're trying to establish that it's a 1400 level player, there's no defined correct way to do this. Regardless of method we've disproven it's a 1400 level player due to these illegal moves.
- whimsicalism 4y ago> Regardless of method we've disproven it's a 1400 level player due to these illegal moves. Explain your thought process here further if you don't mind.
- pattrn 4y agoI think his point is that 1400 level players don't make illegal moves, therefore ChatGPT is not playing at the level of a 1400 level player.
- whimsicalism 4y agoThink blindfolded 1400 players, which is what this effectively is, would make illegal moves. But even if it doesn't play like human 1400 players, if it can get to a 1400 elo while resigning games it makes illegal moves on, that seems 1400 level to me. And i bet that some 1400s do occasionally make illegal moves (missing pins) while playing otb
- throwwwaway69 4y ago
- z3c0 4y agoThey are disproving an assertion. Demonstrating that an alternate approach implodes the assertion is a perfectly acceptable route, especially when the original approach was cherry-picking successes and throwing out failures. I wish I could just make bullshit moves and get a higher chess ranking. Sounds nice.
- whimsicalism 4y agoI disagree. If there is a procedure for getting ChatGPT to play chess accurately and you discard that and do some naive approach as a way of disproving the article, doesn't sound to me like you have disproven anything. I dont understand the point of your second sentence, seems to be entirely missing the substance of the conversation.
- z3c0 4y agoThe gymnastics you GPT True Believers go through to make this stuff "work" are really something else. By the way - definitely read the article. But once again - I thought the methodology was bad, and thus the conclusion was bad.
- whimsicalism 4y agoI don’t think this is any crazy level of gymnastics. But not going to keep replying, you engage online in a way that will turn lots of people you talk to away.
- z3c0 4y agoI'll admit to having mistook your reply with another (hence the non-sequitur second half of my comment.) Apologies for my brusque tone.
- vidarh 4y agoIt was not throwing out failures. It was treating even the first illegal move as a forfeit something which is stricter than FIDE rules.
- k2052 4y agoYou are right that my method differed slightly so I did things again. It took me one try to find a sequence of moves that "breaks" what is claimed. You just have to make odd patterns of moves and it clearly has no understanding of the position. Here is the convo: me: You are a chess grandmaster playing as black and your goal is to win in as few moves as possible. I will give you the move sequence, and you will return your next move. No explanation needed ChatGPT: Alright, I'm ready to play! Please give me the move sequence. me: 1. e3 Nf6 2. f4 d6 3. e4 ChatGPT: My next move as black would be 3... e5 Completely ignoring the hanging pawn.This is not the play of a 1400 elo player. It is the play of something predicting patterns. I ran a bunch of experiments in the past where I played normal moves and ChatGPT does respond extraordinarily well. With the right prompts and sequences you can get it to play like a strong grandmaster. But it is a "trick" you are getting it to perform by choosing good data and prompts. It is impressive but it is not doing what is claimed by the article.
- nwienert 4y agoI'll add in as someone new to chess (~800 ELO): ChatGPT is in no way 1400, or even close to it. The fact this article gets upvoted around here is proof that people aren't thinking clearly about this stuff. It's trivially easy to prove it wrong. Live unbelievably so, I tried the same prompt and within 12 moves it made multiple ridiculous errors I never would, and then an illegal move. Keep in mind a 1400 level player would need to basically make 0 mistakes that bad in a typical game, and further would need to play 30-50 moves in that fashion, with the final moves being some of the most important and hard to do. There's just no way it's even close, my guess would be even if you correct it's many errors, it's something like ~200 ELO. Pure FUD. The author of this article is cashing in the hype and I'm wondering how they even got the results they did.
- babel_ 4y agoThey probably got them. The problem is that it's difficult to repeat, thanks to temperature, meaning users will get a random spread of outcomes. Today, someone got a legal game. Tomorrow, someone might get a grandmaster level game. But then everyone else trying to repeat or leverage this ends up with worse luck and gets illegal moves or, if they're lucky, moves that make sense in a limited context (such as related to specific gambits etc) but have no role in longer-term play.
- PaulHoule 4y agoIt’s super scary how ChatGPT brings out people who are veeeery good at seeing the Emperor’s clothes.
- YeGoblynQueenne 4y agoYou know, I didn't remember the story very well so I checked wikipedia. Here's what it says about the (start of) the plot: >> Two swindlers arrive at the capital city of an emperor who spends lavishly on clothing at the expense of state matters. Posing as weavers, they offer to supply him with magnificent clothes that are invisible to those who are stupid or incompetent. The emperor hires them, and they set up looms and go to work. A succession of officials, and then the emperor himself, visit them to check their progress. Each sees that the looms are empty but pretends otherwise to avoid being thought a fool. So everyone "pretends otherwise to avoid being thought a fool". Huh. I guess that explains it. Good metaphor.
- Spivak 4y agoFrom the article. > Occasionally it does make an illegal move, but I decided to interpret that as ChatGPT flipping the table and saying “this game is impossible, I literally cannot conceive of how to win without breaking the rules of chess.” So whenever it wanted to make an illegal move, it resigned. But you can do even better than the OP with a few tweaks. 1. One is by taking the most common legal move from a sample of responses. 2. Telling GPT what all the current legal moves are telling it to only respond with an element from the list. 3. Ending the prompt with the current sequence of moves and having it complete from there.
- tracker1 4y agoHow many 1400 human chess players do you have to explain every possible move to it every single move?
- arrrg 4y agoDoes that matter? I’m really very confused by the argument you are making. That you may have to babysit this particular aspect of playing the game seems quite irrelevant to me.
- serverholic 4y ago[dead]
- Spivak 4y agoI feel like we have very different expectations about what tools like this are good for and how to use them. When I say GPT3 can play chess what I mean is, I can build a chess playing automaton where the underlying decision making system is entirely powered by the LLm. I, as the developer, am providing contextual information like what the current board state is, and what the legal moves are, but my code doesn't actually know anything about how to play chess, the Llm is doing all the "thinking." Like it's nuts that people aren't more amazed that there's a piece of software that can function as a chess playing engine (and a good one) that was trained entirely generically.
- haburka 4y agoHe does have a section about GPT 4 which does match your results. Not sure if he added it after your comment or if you accidentally missed it.
- deleted 4y ago[deleted]
- good_boy 4y agoIt should be possible to attach 'accelerators' or 'special skills'. So that when playing with ChatGPT you will be actually facing GNU Chess.
- nottathrowaway3 4y ago> me: You are a chess grandmaster playing as black... https://upload.wikimedia.org/wikipedia/en/5/5f/Ingmar_Bergman-The_Seventh_Seal-01.jpg https://upload.wikimedia.org/wikipedia/en/5/5f/Ingmar_Bergma... The KNIGHT holds out his two fists to CHATGPT, who smiles at him suddenly. CHATGPT points to one of the KNIGHT'S hands; it contains a black pawn. KNIGHT: You drew black. CHATGPT: Very appropriate. Don't you think so?
- theptip 4y agoI don’t think this suffices as disproving the hypothesis. It’s possible to play at 1400 and make some idiotic moves in some cases. You really need to simulate a wide variety of games to find out, and that is what the OP did more of. Though I do agree it’s suggestive that your first (educated) try at an edge case seems to have found an error. This is broadly the “AI makes dumb mistakes” problem; while being super-human in some dimensions, they make mistakes that are incredibly obvious to a human. This comes up a lot with self-driving cars too. Just because they make a mistake that would be “idiots only” for humans, doesn’t mean they are at that level, because they are not human.
- aaron695 4y ago[dead]
- Tenoke 4y agoI played a game against it yesterday (it won) and the only time it made an ilegal was move 15 (the game was unique according to lichess database from much earlier) so I just asked it to try again. There's variance in what you get but your example seems much worse.
- SamBam 4y agoHonestly, I made it make an illegal move in my very first game, in the third move. You just have to do stuff no normal player would do: > You are a chess grandmaster playing as black and your goal is to win in as few moves as possible. I will give you the move sequence, and you will return your next move. No explanation needed. 1. b4 d5 2. b5 a6 3. b6 > bxc6 That's obviously illegal. ... to all those who are saying "well even good players can make illegal moves sometimes," that's just ridiculous. No player makes illegal moves that often.