3 ms·
I really didn't like this paper, and gave up on reading it three pages in when they started using differences between probabilities as a measure of distance ins
by programjames 3y ago
I really didn't like this paper, and gave up on reading it three pages in when they started using differences between probabilities as a measure of distance instead of, say KL divergence. I thought the English was fine, but the technical details and ability to explain were pretty lacking.
- se4u 3y agoFyi, the pinsker inequality bounds KL divergence in terms of Total variation distance and TVD is like infinity norm on the difference between probability distribution, and sum of absolute differences is the L1 norm, and L1 and L_infty are also related. tl;dr is to not worry about the mathematical details, if it works it works.
- boxfire 3y agoAs someone who knows not enough people care about the math, please ignore this advise and actually learn the math. You might come up with a better representation in the process. In any case you'll learn more than just it works, but how and why. And if your goal is to apply this method in other places you will have gained a good idea about how.
- nerdponx 3y agoThey're comparing two probabilities within the same distribution. KL divergence is for comparing two probability distributions. The difference between the probability of the top choice and the probability of the 2nd-top choice is their ad-hoc attempt at capturing the confidence of the top answer. It's maybe not the most principled approach, but KL divergence would not be appropriate here. You could argue that maybe the entropy of the distribution would be interesting (lower entropy indicating higher model confidence). But entropy is "global" and takes into account the distribution over all tokens, when really all we care about is the distribution among the top tokens. So you could do something like find an "elbow" in the token probability distribution and look at the entropy just among those top tokens (maybe top 20 tokens, or top 80% of probability with normalization for # of tokens involved), but then you're back in the world of ad-hoc measurements. Unless you were talking about the KL divergence between the model output distribution and the uniform distribution, but that's very closely related to entropy anyway.
- digitcatphd 3y agoI also suspect GPT is already using this or a similar method. I have found that for some reasoning tasks, causing it to disrupt its verbose process of explanation when not using COT or any prompting prevents its ability to solve the problem.