2 ms·
Practically speaking, it's a simple measure of how similar two probability distributions are, minimised (with value zero) when they are the same. So it's often
by JonyEpsilon 3y ago
Practically speaking, it's a simple measure of how similar two probability distributions are, minimised (with value zero) when they are the same. So it's often used as a loss term in optimisations when you want two distributions to be pushed towards being similar. Sometimes this motivated by clever reasoning about information/probability ... but often it's more just "slap a KL on it", because it tends to work.