8 ms·
I felt like I finally understood Shannon entropy when I realized that it's a subjective quantity -- a property of the observer, not the observed. The entropy o
by glial 2y ago
I felt like I finally understood Shannon entropy when I realized that it's a subjective quantity -- a property of the observer, not the observed.
The entropy of a variable X is the amount of information required to drive the observer's uncertainty about the value of X to zero. As a correlate, your uncertainty and mine about the value of the same variable X could be different. This is trivially true, as we could each have received different information that about X. H(X) should be H_{observer}(X), or even better, H_{observer, time}(X).
As clear as Shannon's work is in other respects, he glosses over this.
- JumpCrisscross 2y ago> it's a subjective quantity -- a property of the observer, not the observed Shannon's entropy is a property of the source-channel-receiver system.
- glial 2y agoCan you explain this in more detail? Entropy is calculated as a function of a probability distribution over possible messages or symbols. The sender might have a distribution P over possible symbols, and the receiver might have another distribution Q over possible symbols. Then the "true" distribution over possible symbols might be another distribution yet, call it R. The mismatch between these is what leads to various inefficiencies in coding, decoding, etc [1]. But both P and Q are beliefs about R -- that is, they are properties of observers. [1] https://en.wikipedia.org/wiki/Kullback–Leibler_divergence#Coding https://en.wikipedia.org/wiki/Kullback–Leibler_divergence#Co...
- rachofsunshine 2y agoThis doesn't really make entropy itself observer dependent. (Shannon) entropy is a property of a distribution. It's just that when you're measuring different observers' beliefs, you're looking at different distributions (which can have different entropies the same way they can have different means, variances, etc).
- mitthrowaway2 2y agoEntropy is a property of a distribution, but since math does sometimes get applied, we also attach distributions to things (eg. the entropy of a random number generator, the entropy of a gas...). Then when we talk about the entropy of those things, those entropies are indeed subjective, because different subjects will attach different probability distributions to that system depending on their information about that system.
- stergios 2y ago"Entropy is a property of matter that measures the degree of randomization or disorder at the microscopic level", at least when considering the second law.
- mitthrowaway2 2y agoRight, but the very interesting thing is it turns out that what's random to me might not be random to you! And the reason that "microscopic" is included is because that's a shorthand for "information you probably don't have about a system, because your eyes aren't that good, or even if they are, your brain ignored the fine details anyway."
- canjobear 2y agoSome probability distributions are objective. The probability that my random number generator gives me a certain number is given by a certain formula. Describing it with another distribution would be wrong. Another example, if you have an electron in a superposition of half spin-up and half spin-down, then the probability to measure up is objectively 50%. Another example, GPT-2 is a probability distribution on sequences of integers. You can download this probability distribution. It doesn't represent anyone's beliefs. The distribution has a certain entropy. That entropy is an objective property of the distribution.
- mitthrowaway2 2y agoOf those, the quantum superposition is the only one that has a chance at being considered objective, and it's still only "objective" in the sense that (as far as we know) your description provided as much information as anyone can possibly have about it, so nobody can have a more-informed opinion and all subjects agree. The others are both partial-information problems which are very sensitive to knowing certain hidden-state information. Your random number generator gives you a number that you didn't expect, and for which a formula describes your best guess based on available incomplete information, but the computer program that generated knew which one to choose and it would not have picked any other. Anyone who knew the hidden state of the RNG would also have assigned a different probability to that number being chosen.
- dist-epoch 2y agoTrivial example: if you know the seed of a pseudo-random number generator, a sequence generated by it has very low entropy. But if you don't know the seed, the entropy is very high.
- rustcleaner 2y agoTheoretically, it's still only the entropy of the sneed-space + time-space it could have been running in, right?
- sva_ 2y agohttps://archive.is/9vnVq https://archive.is/9vnVq
- canjobear 2y agoWhat's often lost in the discussions about whether entropy is subjective or objective is that, if you dig a little deeper, information theory gives you powerful tools for relating the objective and the subjective. Consider cross entropy of two distributions H[p, q] = -Σ p_i log q_i. For example maybe p is the real frequency distribution over outcomes from rolling some dice, and q is your belief distribution. You can see the p_i as representing the objective probabilities (sampled by actually rolling the dice) and the q_i as your subjective probabilities. The cross entropy is measuring something like how surprised you are on average when you observe an outcome. The interesting thing is that H[p, p] <= H[p, q], which means that if your belief distribution is wrong, your cross entropy will be higher than it would be if you had the right beliefs, q=p. This is guaranteed by the concavity of the logarithm. This gives you a way to compare beliefs: whichever q gets the lowest H[p,q] is closer to the truth. You can even break cross entropy into two parts, corresponding to two kinds of uncertainty: H[p, q] = H[p] + D[q||p]. The first term is the entropy of p and it is the aleatoric uncertainty, the inherent randomness in the phenomenon you are trying to model. The second term is KL divergence and it tells you how much additional uncertainty you have as the result of having wrong beliefs, which you could call epistemic uncertainty.
- bubblyworld 2y agoThanks, that's an interesting perspective. It also highlights one of the weak points in the concept, I think, which is that this is only a tool for updating beliefs to the extent that the underlying probability space ("ontology" in this analogy) can actually "model" the phenomenon correctly! It doesn't seem to shed much light on when or how you could update the underlying probability space itself (or when to change your ontology in the belief setting).
- bsmith 2y agoCouldn't you just add a control (PID/Kalman filter/etc) to coverage on a stability of some local "most" truth?
- bubblyworld 2y ago
- vinnyvichy 2y agoBaez has a video (accompanying, imho), with slides https://m.youtube.com/watch?v=5phJVSWdWg4&t=17m https://m.youtube.com/watch?v=5phJVSWdWg4&t=17m He illustrates the derivation of Shannon entropy with pictures of trees
- IIAOPSW 2y agoTo shorten this for you with my own (identical) understanding: "entropy is just the name for the bits you don't have". Entropy + Information = Total bits in a complete description.
- deleted 2y ago[deleted]
- CamperBob2 2y agoIt's an objective quantity, but you have to be very precise in stating what the quantity describes. Unbroken egg? Low entropy. There's only one way the egg can exist in an unbroken state, and that's it. You could represent the state of the egg with a single bit. Broken egg? High entropy. There are an arbitrarily-large number of ways that the pieces of a broken egg could land. A list of the locations and orientations of each piece of the broken egg, sorted by latitude, longitude, and compass bearing? Low entropy again; for any given instance of a broken egg, there's only one way that list can be written. Zip up the list you made? High entropy again; the data in the .zip file is effectively random, and cannot be compressed significantly further. Until you unzip it again... Likewise, if you had to transmit the (uncompressed) list over a bandwidth-limited channel. The person receiving the data can make no assumptions about its contents, so it might as well be random even though it has structure. Its entropy is effectively high again.
- kragen 2y agoshannon entropy is subjective for bayesians and objective for frequentists
- marcosdumay 2y agoThe entropy is objective if you completely define the communication channel, and subjective if you weave the definition away.
- kragen 2y agothe subjectivity doesn't stem from the definition of the channel but from the model of the information source. what's the prior probability that you intended to say 'weave', for example? that depends on which model of your mind we are using. frequentists argue that there is an objectively correct model of your mind we should always use, and bayesians argue that it depends on our prior knowledge about your mind
- kragen 2y ago(i mean, your information about what the channel does is also potentially incomplete, so the same divergence in definitions could arise there too, but the subjectivity doesn't just stem from the definition of the channel; and shannon entropy is a property that can be imputed to a source independent of any channel)
- marcosdumay 2y ago> he glosses over this All of information theory is relative to the channel. This bit is well communicated. What he glosses over is the definition of "channel", since it's obvious for electromagnetic communications.