3 ms·
> Since there are ~5000 dimensions, in which of those dimensions are we moving k out ? > Is the idea you just move, out, in all dimensions such that the final
by mhink 3y ago
> Since there are ~5000 dimensions, in which of those dimensions are we moving k out ?
> Is the idea you just move, out, in all dimensions such that the final Euclidean distance is k ?
If I understand correctly, the basic idea is that in earlier experiments, they were able to use a relatively-simple technique to come up with what they call a "probe vector" in the embedding space which represents "the property of a word starting with <letter>". For any given token, the authors established that it was 98% probable that the token's embedding vector would be closer (by cosine similarity) to the "probe vector" representing the first letter of that word than any other probe vector. This is shown in the first graph of the section "A puzzling discovery".
With that in mind, the diagram below that should start making more sense: "emb" is a particular token's embedding vector and "probe" is the probe vector for the token's first letter. "emb_proj" is the projection of "emb" onto "probe".
What they're doing is tweaking the network weights by subtracting multiples of `emb_proj` from `emb` (where the specific multiple is the parameter K), and then seeing how it behaves differently for different values of K.
Their original observation when doing this was that it reliably caused the model to claim that the first letter of the tweaked word was not the letter in question. In this article, they're trying to figure out how far they can push the tweak and still get reasonably accurate definitions of a token.
What they discovered is that when they push a token's embedding vector further and further out along its "first-letter vector" and ask the network to define that word, the definitions it provides seem to follow particular themes during different regimes of K.