3 ms·
Slightly off-topic, but I am just starting to learn about the field, and saw > GPT4 still projects information into some highly non-linear latent space and sam
by JoshuaDavid 3y ago
Slightly off-topic, but I am just starting to learn about the field, and saw
> GPT4 still projects information into some highly non-linear latent space and samples from that
Do you have more information/know where I can find things to read about the "highly non-linear" part of that? I've been reading some stuff[1][2][3] about smaller models, and my impression has been that the latent space is shockingly linear for those ones. But counterexamples would be informative, especially if it's something that only happens with larger models.
-----
[1] https://arxiv.org/abs/2209.15162 https://arxiv.org/abs/2209.15162
[2] https://www.neelnanda.io/mechanistic-interpretability/othello https://www.neelnanda.io/mechanistic-interpretability/othell...
[3] https://arxiv.org/pdf/1610.01644.pdf https://arxiv.org/pdf/1610.01644.pdf
- jiggawatts 3y agoHilariously, the best explanations for how AI models like GPT work I got from GPT-4 itself! There’s a lot of jargon and assumed shared knowledge in research papers that make them hard for a beginner to parse. I’ve found GPT can “translate” and explain in context. I know neither Python nor PyTorch, so I give it snippets and tell it to show me the Mathematica equivalent, which I can read.
- psyklic 3y agoThe space must be non-linear, as a consequence of the non-linear activation functions. A neural net is just a big math equation, and deeply embedded throughout are non-linear transforms which necessarily make the entire transform non-linear. Just like the presence of 1/x makes an equation no longer linear (at least with rare exception!). Here is a deeper explanation/visualization: https://colah.github.io/posts/2014-03-NN-Manifolds-Topology/ https://colah.github.io/posts/2014-03-NN-Manifolds-Topology/ The three references don't seem to conclude that the latent space is linear. The first seems to mention "linear" only since they add an additional linear layer to project an image encoding into a generative language model. So I'm not sure this applies here, since the encoding itself is already richly complex by the time it is mapped. The second and third are using a "linear probe" in an attempt to gain insights about each layer. This feeds the output of a given layer into a linear classifier that attempts to predict the correct output labels. This doesn't perform well in early layers, but it improves monotonically until the final layer is reached. The researchers conclude this happens entirely as a consequence of the final layer being a linear classifier. So, the features eventually must become linearly separable since that's what the network was trained to do. This doesn't conclude that each layer's feature space is linear. Instead, they are just using a linear projection to examine how "easily" the net at that layer can make correct predictions. Even if a layer's output is decently predictive in this way, the actual representation could still be richer and contain additional information.
- JoshuaDavid 3y ago> So, the features eventually must become linearly separable since that's what the network was trained to do. This clarifies things a lot, thanks!