3 ms·
Great article. It was a joy to read. I have one question though: Why do we integrate the control vector across all layers of a neural network, rather than limit
by WiSaGaN 3y ago
Great article. It was a joy to read. I have one question though: Why do we integrate the control vector across all layers of a neural network, rather than limiting its application to just the final layer or a subset of layers? Given that each vector influences every layer it passes through, resulting in a cumulative effect, isn't there a risk of excessively skewing the data representation?
- semi-extrinsic 3y agoAs the author stated in this post, it's not actually one vector, but a list of one vector per layer. If I understand it correctly, these vectors can have different total magnitude across the layers. If the PCA (or other technique) identifies that layers 17, 36 and 41 are important for "concept X", the vectors for those layers will be the strongest when repeng'ing for that concept.
- danm1618 3y agoA note worth mentioning is that the PCA is not trained across the layers, but independently on each layer, across the provided examples. Nevertheless, it's conceivable that specific layers could possess a significant control vector, but not solely because of directly leveraging the first principal component.
- sigmoid10 3y agoThe final layer will not encode high level concepts anymore, it's essentially just tokens from the vocabulary. It would be impossible to encode abstract things like "niceness" in it. As long as we don't know exactly at which layers this behaviour emerges, randomly choosing a subset also won't work. So what they did is apply a custom vector to every layer and let PCA figure out which of these vectors are actually necessary. Curiously, looking at these vectors should also tell you more about where and how the model processes these things.