3 ms·
Have a look at this research paper[1] and follow-up post[2]. The upshot is that the researchers trained a model of the same shape that the main language models
by JoshuaDavid 4y ago
Have a look at this research paper[1] and follow-up post[2]. The upshot is that the researchers trained a model of the same shape that the main language models are to play Othello. The model started with no knowledge of the game, and was fed a bunch of move sequences (e.g. "c4 c3 d3 e3"). The researchers were then able to take that model, and figure out, by looking only at the activations, what the board state was. Changing the activations so that they modeled a different board state caused the model to play moves that made sense with the new board state, and did not make sense with the original board state.
And then the follow-up research[2] established that, not only can you determine the board state by looking at the activations, but that you can trivially do so. Specifically, if you look at the residual stream after layer 4, you can build a linear classifier for whether each of the 64 squares is empty, and whether each of the 64 squares contains a token owned by the player whose turn it is. If you bump the values in the direction implied by each of those classifiers, you can make the model output moves in the same way as if the tokens on those squares had the opposite color. There's even a colab notebook[3] you can play with.
That research was on Othello, not chess, but I'm pretty sure that "LLMs are able to develop models of the world if doing so helps them predict the next token better" generalizes to chess too.
[1] https://arxiv.org/pdf/2210.13382.pdf https://arxiv.org/pdf/2210.13382.pdf
[2] https://www.neelnanda.io/mechanistic-interpretability/othello https://www.neelnanda.io/mechanistic-interpretability/othell...
[3] https://colab.research.google.com/github/likenneth/othello_world/blob/master/Othello_GPT_Circuits.ipynb https://colab.research.google.com/github/likenneth/othello_w...