3 ms·
After the phrase:Manipulating the logarithms, we can also get ... the formula is incorrect, since p_j have disappeared. \[D_{KL}(P,Q)=-\sum_{j=1}^{n}\log_2 \fr
by meanpp 1y ago
After the phrase:Manipulating the logarithms, we can also get ... the formula is incorrect, since p_j have disappeared.
\[D_{KL}(P,Q)=-\sum_{j=1}^{n}\log_2 \frac{q_j}{p_j}=\sum_{j=1}^{n}\log_2 \frac{p_j}{q_j}\]
The post is just basic definitions and simple examples for cross entropy and KL divergence.
There is a section about the relation of cross entropy and maximum likelihood estimation at the end that seems not so easy to understand but implies that the limit of a estimator applied to a sample from a distribution is the KL divergence when the sample length tends to infinity.