3 ms·
At a high level: It's an article observing how GPT-J token embeddings are positioned in 'embedding space', with some connections drawn to GPT 3. They then exp
by netruk44 3y ago
At a high level:
It's an article observing how GPT-J token embeddings are positioned in 'embedding space', with some connections drawn to GPT 3.
They then experiment with having GPT-J provide definitions for "nokens" (A "noken" is basically a made up embedding created by modifying the embedding generated from a real token) to see what happens when a model is presented with a novel embedding outside of the trained embedding space to see how it interprets them.
Diving in a little more:
An observation from the article is that the first letter of a word represented by an embedding is actually encoded into that embedding such that you could identify what letter any embedding starts with, with 98% certainty.
They observe this linear relationship and devise an experiment. They ask GPT-J what letter the word "icon" starts with, and it correctly replies "I". They then create a "noken" for the word "icon" and modify it so that it no longer represents that the word starts with "I" and ask GPT-J again what letter this "noken" starts with. GPT-J then incorrectly replies that the word "icon" does not start with the letter "I".
So then they pose the question "can the model still define the word 'broccoli' even if we shift the first letter away from 'B'?" and the answer is almost always yes, changing the semantic "first letter" of an embedding doesn't change the model's ability to understand the embedding.
They then do a series of experiments asking GPT-J to provide the "typical definition for the word '<noken>'." with the example word 'hate' being used. They gradually shifted the first letter away from 'H', and see that up until very high levels of modification, the model can still give a normal definition for the modified 'hate' noken ("a strong feeling of dislike or hostility.").
At extreme modifications to the noken, the model starts providing strange definitions ("a person who is not a member of a particular group.", "a period of time during which a person or thing is in a state of being").
They repeat the experiment and note that this behavior of the definition keeping stable up until 'collapse' is common across many tokens:
> ...usually ending up with something about a person who isn't a member of a group by k = 100, having passed through one or more other themes involving things like Royal families, places of refuge, small round holes and yellowish-white things, to name a few of the baffling tropes that began to appear regularly.