3 ms·
Negative log is the only decreasing function which satisfies f(x * y) = f(x) + f(y). If you have two independent events it makes sense that the total surprisal
by eutectic 3y ago
Negative log is the only decreasing function which satisfies f(x * y) = f(x) + f(y). If you have two independent events it makes sense that the total surprisal (information content) should be the sum of the surprisals of the two events.
Another way to see it is as a continuous generalization of the idea that you need n bits to represent 2^n equally likely alternatives.
- beyondCritics 3y ago> Negative log is the only decreasing function which satisfies f(x * y) = f(x) + f(y) Times a constant!
- sarma17 3y agothanks for the reply but I don't think I understood what you said in the context of my question. I'm stuck on why we care about surprisal as `log 1/p`.
- sixo 3y agoI think "information" is a better name than surprisal. If your distribution has N equally-likely values, `p(x) = 1/N`, and information/surprisal `I(x) = log(N)`. In base 2, this is how many bits are required to specify exactly WHICH of the N values you're talking about. If `x` is not a single one of the N states but an event consisting of `n(x)` states, then `I(x) = log(N) - log(n(x))`, suggesting it takes _somewhat less information_ to specify this particular state, and it correctly gives 0 if `n(x) = N`, i.e. there's only one state. Exactly what this "less information" means is vague, but you might think of it in terms of compression: if you compress some stream of data which is sampled from `X` with probability `p(x)`, you could use use the shortest codes (0, 10, 11, etc) for the most common values with some "stop word" to say when the end of a datum is reached. `I(x)` captures this sense in general, but it might only become literally true in the limit of a very large stream of data with a very large dictionary.
- eutectic 3y agolog 1/p is just -log p
- ramblenode 3y ago> If you have two independent events it makes sense that the total surprisal (information content) should be the sum of the surprisals of the two events. I think the parent is also asking why we would expect surprisal to be additive rather than multiplicative like probabilities.
- Majromax 3y agoBecause if two things happen – one totally expected and one very surprising – then on net you’re still surprised.
- ramblenode 3y agoThat's also the case for probabilities.
- joelthelion 3y agoI think in part because it's really nice to measure information in bits rather than tiny probabilities. And it aligns well with how information is stored.
- KRAKRISMOTT 3y agoConvolution?