7 ms·
Question here is why gpt-3.5-instruct can then beat stockfish.
by computerex 2y ago
Question here is why gpt-3.5-instruct can then beat stockfish.
- fsndz 2y agoPS: I ran and as suspected got-3.5-turbo-instruct does not beat stockfish, it is not even close "Final Results: gpt-3.5-turbo-instruct: Wins=0, Losses=6, Draws=0, Rating=1500.00 stockfish: Wins=6, Losses=0, Draws=0, Rating=1500.00" https://www.loom.com/share/870ea03197b3471eaf7e26e9b17e1754?sid=073f0b78-7be9-42af-95c2-240d3b49d1c8 https://www.loom.com/share/870ea03197b3471eaf7e26e9b17e1754?...
- computerex 2y agoMaybe there's some difference in the setup because the OP reports that the model beats stockfish (how they had it configured) every single game.
- Filligree 2y agoOP had stockfish at its weakest preset.
- fsndz 2y agoDid the same and gpt-3.5-turbo-instruct still lost all the games. maybe a diff in stockfish version ? I am using stockfish 16
- mannykannot 2y agoThat is a very pertinent question, especially if Stockfish has been used to generate training data.
- golol 2y agoYou have to get the model to think in PGN data. It's crucial to use the exact PGN format it sae in its training data and to give it few shot examples.
- bluGill 2y agoThe artical appears to have only run stockfish at low levels. you don't have to be very good to beat it
- lukan 2y agoCheating (using a internal chess engine) would be the obvious reason to me.
- TZubiri 2y agoNope. Calls by api don't use functions calls.
- permo-w 2y agothat you know of
- TZubiri 2y agoSure. It's not hard to verify, in the user ui, function calls are very transparent. And in the api, all of the common features like maths and search are just not there. You can implement them yourself. You can compare with self hosted models like llama and the performance is quite similar. You can also jailbreak and get shell into the container to get some further proof
- permo-w 2y agothis is all just guesswork. it's a black box. you have no idea what post-processing they're doing on their end
- girvo 2y agoHow can you prove this when talking about someones internal closed API?
- nske 2y agoBut in that case there shouldn't be any invalid moves, ever. Another tester found gpt-3.5-turbo-instruct to be suggesting at least one illegal move in 16% of the games (source: https://blog.mathieuacher.com/GPTsChessEloRatingLegalMoves/ https://blog.mathieuacher.com/GPTsChessEloRatingLegalMoves/ )
- shric 2y agoI'm actually surprised any of them manage to make legal moves throughout the game once out of book moves.