4 ms·
It all depends on what prompt you use though. You can just tell all current frontier models to write a chess engine first, and then play a game of chess against
by tzone 17d ago
It all depends on what prompt you use though. You can just tell all current frontier models to write a chess engine first, and then play a game of chess against you using that engine.
It will probably do a pretty good job if you ask it that way (it will also burn a shit ton of tokens, but hey, that is part of the fun).
On that note, I actually had an overall harness (for experimenting) that was essentially like this:
"for any task, instead of answering question directly, write a program to answer the question instead. test and verify the program before giving the answer".
It actually worked incredibly well on all "gotcha" LLM questions like math or counting letters in words and all sorts of stuff.
Of course it was ridiculously slow and very expensive but it was a proof of concept that it can actually be much more accurate on every task if you are willing to spend infinite amount of money.