3 ms·
The K-L divergence is relevant there, even though I'm pretty sure that that "formally" comment is meant as a joke and not serious. (wikipedia) "A simple interp
by phiresky 4y ago
The K-L divergence is relevant there, even though I'm pretty sure that that "formally" comment is meant as a joke and not serious.
(wikipedia) "A simple interpretation of the KL divergence of P from Q is the expected excess surprise from using Q as a model when the actual distribution is P.
The sentence you quoted posits that whenever the LMM is in a state where it is "simulating" a nice and helpful person, the simulation is also consistent with an insane, violent person that's currently pretending to be nice, but not the other way around.
The author isn't talking about the loss or error of predicting individual tokens. If you look the larger scale behaviour, predicting, e.g., a scalar niceness value of the response based on one of two "modes" that you assume the LLM is currently in (either Waluigi or Luigi), then you'll be less surprised if Waluigi acts like Luigi than the other way around.
The probability distribution of niceness when assuming the LLM is in a state of "Luigi" would have a high mean and low variance, while the distribution for Waluigi would have a lower mean but a higher variance.
Thus, the KL divergence of Waluigi (that is, the probability distribution of niceness you'd predict when assuming the model is in Waluigi mode) from Luigi would be high, while the other way around `KL(Luigi, Waluigi)` would be low.
It should be easy to construct an example with concrete values using two normal probablility distributions.