11 ms·
Since this topic isn't so well-known, I wrote the case arguing that frequentist interpretations don't work, but algorithmic information theory (Kolmogorov compl
by EbTech 6y ago
Since this topic isn't so well-known, I wrote the case arguing that frequentist interpretations don't work, but algorithmic information theory (Kolmogorov complexity) does. I want to make this accessible and persuasive, so thoughts, questions, and arguments would be appreciated!
- opheliate 6y agoSuch an interesting post, thank you for sharing! I'm in my second year of a maths degree currently, and we obviously studied frequentist probability/stats in the first year, but I'm not taking any probability modules this year. I found the tone & accessibility was just right for me :)
- enriquto 6y agoA simple sentence that I've found useful for pedagogy: "the probability of that coin toss being 50% does not talk about the coin; it talks about you, and about your partial knowledge of the universe." You can add: "The coin toss itself is deterministic and the result can be computed if you know the initial position and speed." They will inevitably bother you about the physical impossibility to measure the starting position and speed exactly, and then you say "ok, forget about the coin. You have 5 white and 5 black balls inside this opaque cylinder. What's the probability that the top ball is white? This does not talk about the balls (the color of the top one is already determined) but about your partial knowledge of them". (EDIT: formatting)
- alisonkisk 6y agoHow does the ball situation help? It has the same problems of physical impossibility of measuring however the balls were ordered. (Modelling someone's brain?) I guess the argument works on someone with an unscientific model of the human brain, but that's one step forward and two steps back.
- enriquto 6y agoSomebody just put the balls there carefully, and did not tell you in what order, just how many of each color.
- Frost1x 6y agoThat still boils down to a lack of prior information which I don't think removes the argument for "I don't have enough information." Probably have to use actual quantum phenomena that behave probabilistically by definition if you want a currently irrefutable physical example. I'm personally not convinced even this is fundamentally probabilistic and we currently have to rely on probability theory as a crutch for complex behaviors we just quite don't understand yet or don't have the time and resources to compute.
- posterboy 6y agoYou don't need quantum physics to formulate a philosophical standpoint that happens to agree with the Kopenhagen Interpretation. I'm not sure if it helps, but I suppose a compromise here would be the assumption that you don't really know the starting configuration of yourself, why you draw probabilistic inferences naturally, that the sun will go up tomorrow like every day. If that has a biologic explanation, then the top comment was not just to the illusive argument of platonic ideals.
- enriquto 6y ago> "I don't have enough information." My point exactly. Probability theory is a precise mathematical formalization of the concept of "not enough information".
- st1x7 6y ago> the probability of that coin toss being 50% does not talk about the coin; it talks about you, and about your partial knowledge of the universe. But it does talk about the coin - a weighted coin would have a different probability. Same in the example with the white/black balls - if they weren't 5 white and 5 black but 6 white and 4 black, the probability you would assign to the top one would be different. Again, the probability is a way to describe the balls themselves, not just our knowledge. I get the general idea of representing probability as uncertainty and partial knowledge but your statements strike me as just straight up incorrect.
- wearsshoes 6y agoThis is still a partial knowledge situation - you have the information that there are a certain proportion of colored balls in the chamber, but not the information about their order. The probability includes the information we do know and allows inferences about information we don’t know.
- st1x7 6y agoSure, I just can't agree with the parent statement that probability has nothing to do with the object it describes, which is demonstrably false.
- StavrosK 6y agoI don't think it's demonstrably false: If you don't know that the coin is weighted, the probability is 50%. Probabilities are predictions and estimates, not fundamentally about the thing itself, but about what we know about the thing.
- deleted 6y ago[deleted]
- mannykannot 6y agoThe principle of indifference? I know that it is a commonplace assumption, but feels to me as though one is assuming one has more information than is justified. Coming back to the article's "economist's wager", is it rational to bet with even odds on something you know nothing about? If the assumption is interpreted as a testable hypothesis about outcomes, why would complete ignorance imply any particular result? On the other hand, if it is interpreted strictly as a statement about one's knowledge, why present it exactly as if one had sufficient knowledge of the situation to know that the probability is 0.5? Maybe the author will have an answer in part 2.
- Sharlin 6y agoThis is basically the Bayesian interpretation of probability.
- enriquto 6y agoOf course. But if you pronounce a fancy word like "bayesian" there's a large amount of minds that shut irremediably.
- v64 6y agoThat's also why we call it QBism [1] instead of quantum bayesianism [1] https://en.wikipedia.org/wiki/Quantum_Bayesianism https://en.wikipedia.org/wiki/Quantum_Bayesianism
- lottin 6y agoSaying that "the probability of a coin toss of 50% talks about you" is not an interpretation of probability. Saying that we are "50% sure" is also not an interpretation of probability. It's a nonsensical statement. It's like saying we are "50% angry". It doesn't really mean anything.
- alisonkisk 6y agoI don't understand your claims that these statements are are meaningless. They are commonly uttered and understood.
- lottin 6y agoI can understand expressions such as "pretty sure" or "completely sure". I do not understand the expression "to be X% sure". If someone says they're "37% sure" tomorrow will rain, what does that mean exactly?
- dinosaurdynasty 6y ago37% of the time that someone says they are 37% sure of a statement X the statement X is true (assuming they're calibrated correctly/etc).
- nimbleal 6y agoIt may be deterministic, but are you sure it would be computable? One does not necessarily imply the other.
- kps 6y agoForget the coin and the balls — does the nucleus decay? You're not missing any knowledge; there isn't any.
- jbay808 6y agoThis is definitely the most interesting example, but it's not obvious that a situation where the relevant information is fundamentally inaccessible is a situation where you aren't missing any information. It's your best bet for a scenario where you can be sure that nobody else has more information than you do, though.
- mannykannot 6y agoWigner's friend might have something to say about this... I don't have a specific argument to make here, only the feeling that if it were all just a matter of what a given observer knows, no-one would be talking about there being a QM measurement problem.
- codethief 6y agoThis is a very good point! In fact, some people do argue that there is no measurement problem in the Copenhagen formulation of quantum mechanics to begin with – at least if you take it seriously and strictly go by the rule that the laws laid down by Bohr et al. only concern you as the observer and your knowledge about the system, and not the system itself. Following this train of thought, there is nothing "real" about the wavefunction and it is just a tool to come up with predictions. The same goes for the collapse of the wave function (which just describes a change in your ability to predict future measurements, and not a change of the object) and the term "measurement" (which we might as well replace with "enlightenment", i.e. the moment in which we obtain knowledge about the system). In that sense, the only difference between classical and quantum mechanics is that our knowledge (viewed as a mathematical quantity) behaves differently in both theories: In classical physics, when we conduct multiple measurements of a given system in a row, our knowledge about that system will increase – to the point that, once we have measured all system properties to sufficient accuracy, we'll able to predict what any future measurement of any of those properties will yield (again, with some predictable uncertainty). So the knowledge of all our measurements has added up, it is an additive quantity. In QM, this is fundamentally different: We can only know anything about the object the very moment we look at it. The rules of quantum mechanics (again, in the very strict interpretation laid out above) dictate that the second we conduct a measurement, we can forget about any knowledge obtained through previous measurements of other (conjugate) observables: Future measurements of those observables are inherently unpredictable. In that sense, our knowledge about quantum-mechanical objects never "adds up" to anything. (To see that this is really the the distinguishing feature between classical and quantum mechanics, recall that the existence of conjugate observables really is the only thing setting apart the quantum from the classical world: Without conjugate observables it would be impossible to distinguish, say, 100 electrons in a superposition of spin up and down from an ensemble of 100 electrons of which 50 are in a spin up state and the other 50 are in a spin down state.) Of course, this whole interpretation is very unsatisfactory to lots of people (myself included) for a whole bunch of reasons. I assume that, to a large degree, this is due to the fact that laws of nature that put human observers in their very center seem rather undesirable. (At least since the time we switched from a geocentric to a heliocentric view of the world.) But my impression is that there's another reason: Our intuition from classical mechanics & statistics has taught us that objects exist independently of us as observers and behave in a deterministic fashion, at least provided we as observers know enough about them. (Meaning that the more we know about the coin's initial position and velocity, the more likely we are to predict the outcome of the coin toss. If we don't know anything about the coin, though, the outcome is as unpredictable as measuring spin up/down in quantum mechanics.) Unfortunately, this whole line of argument is circular: The reason we believe that the existence of physical objects is independent of us, is precisely because knowledge in classical mechanics is an additive quantity and we can get to the point where we know "enough" to come up with deterministic predictions. That is, we never have to discard knowledge when running new measurements and so our knowledge takes on a independent "role" – which we call reality.
- MereInterest 6y agoAs a caveat, while this intuition works for classical mechanics, it does not work for quantum mechanics. All observations are consistent with wave function collapse being fundamentally random. Any hidden variables would need to be transmitted many times faster than the speed of light (~10000x, last time I checked the experiments), and are therefore inconsistent with our understanding of special relativity.
- smallnamespace 6y agoNote that pilot wave theory, an (out of vogue) interpretation of quantum mechanics, also recasts the apparent randomness in quantum mechanics as due to our ignorance of the exact state of the pilot wave. Even Einstein struggled with quantum mechanics, famously saying "[God] does not play dice with the universe".
- JamisonM 6y agoThere seems to be a lot of quibbling about the simple sentence here but I find it clarifying. Discussing coin tosses is a thought experiment with a very practical physical analog so spelling out clearly what the thought experiment's actual subject is has valuable properties so you don't get lost in the weeds of the physical execution of flipping coins.
- deleted 6y ago[deleted]
- alisonkisk 6y agoWhy is frequentism bad because it only gives certainty for infinite samples, but complexity is good despite being non-computable? It's two sides of the same coin -- computable uncertainty va non-computable certainty.
- EbTech 6y agoI don't think frequentism is "bad"; just insufficient as a gold standard interpretation of probabilistic claims. I liked an analogy from the reference by Rathmanner & Hutter: the most "correct" chess-playing program involves a complete search along the tree of possible games. In practice, we try to approximate this ideal. In the case of Kolmogorov complexity, a reasonable takeaway might be to use the shortest program that we're able to find, even if it's not the shortest overall.
- TheOtherHobbes 6y agoOK - so apply Kolmogorov complexity to election polling. How does that work out? I think you're confusing various possible maps with the territory in a less than useful way. Given that frequentist interpretations are approximations - and understood as such - and Kolmogorov complexity isn't computable at all, what problem have you solved here?
- EbTech 6y agoHm I admit it's hard to talk convincingly about election prediction, since we don't have practical algorithms to do this; a lot of it comes down to human judgment. The philosophical point (which might be approximated algorithmically someday, or by intelligent minds today) is that your election probabilities should come out of an overall highly compressed model of the world. In theory, a Bayesian who uses the prior 2^-K(x) over all strings x should, with sufficient life experience, come up with good estimates, in a certain sense. I'll have to think about this example more carefully when fulfilling my promise of writing about how this theory relates to everyday decision-making. Thanks for pointing out a potential weakness :)
- qsort 6y agoThe article is excellent, congratulations. A couple observations/questions. 1) You didn't comment on the bayesian viewpoint that probability reflects a subjective idea about the state of the world. One might argue, for example, that probability isn't measurable, and that therefore, strictly speaking, a statement about the objective probability of an event isn't meaningful. Experimental evaluation would have to be done on an entire model instead. Do you have any objections to that point of view? 2) I don't find the case about Kolmogorov complexity to be actually convincing, at least not as per the requirements the rest of the article sets. "3141592..." could pass as either "random digits" or "first digits of pi". The fact that it's highly unlikely a true RNG would have generated exactly those, we are back to a frequentist argument there. It's likely I'm missing something, could you elaborate more or give me a pointer?
- SmooL 6y agoIsn't that what the occam's razor argument was for? Sure, both an RNG and the "40 digits of pie" program can produce that output, but the "40 digits of pie" program is shorter, therefore having less kolmogorov complexity
- qsort 6y agoYeah but is it? They are both programs with logarithmic Kolmogorov complexity, and Kolmogorov complexity is only defined up to constant factors unless you commit to a computational model. If you do commit to one, which is the shortest is just a function of what are the specifics of the model you chose, which isn't really interesting, it's literally code golf at that point.
- FartyMcFarter 6y ago> They are both programs with logarithmic Kolmogorov complexity In the case of true random numbers, how is that so? Very few random sequences can be generated by a logarithmic-sized program, since most strings are not significantly compressible [1]. A simple counting argument shows that: there are 2^n strings of n bits, but only 2^(lg n) = n logarithmic-sized strings, a much smaller number! [1] http://theory.stanford.edu/~trevisan/cs154-12/kolcomplexity-rev.pdf http://theory.stanford.edu/~trevisan/cs154-12/kolcomplexity-...
- JamisonM 6y agoSome thoughts on the composition: * "I wrote the case arguing that frequentist interpretations don't work, but algorithmic information theory does": if that's what you are up to here then I think it would be for readers if you stated that up front in some way. And hit me with some kind of summary at the end that makes the concise version of your argument at the end, it's a long article. * Shorter might be better: There's a lot of stuff in here that I think you can pare out in the probability discussion that maybe isn't adding that much to your argument. I think there is a lot to be gained by assuming a generous reader. * Betting might be a distraction to your point: This might be confusing the imperfect knowledge of participants in a market with the imperfect knowledge of all the physical forces involved in a physical phenomenon and how that related to the seeming "randomness" of a coin flip for your reader. (The liquidity and stuff.. this is just not related to your point.) * Don't undermine your point with unrelated assumptions: "I imagine they wouldn’t consider their world unlikely at all: they would just add a new law to their description of physics: all dice, as if by divine intervention, are deemed to exhibit this strange behaviour" this lead me to think that you were just sort of shooing away the whole last X decades of high vs. low energy physics, we collectively certainly don't think that we have the rules correct precisely because of this complication, we find the idea that we need 2 sets of rules improbable and believe that there must be a way to explain everything with a single set of rules. So your mythical dice society probably would consider their dice exception a very unlikely world.. they would be confident they have the world wrong! A dubious assertion (or at least one that would need a whole lot of explanation) can be an off-ramp for a subset of readers.
- EbTech 6y agoThanks for the detailed critique! I'll take some time to think about how to better make the points that I wanted to convey with those sections.
- richard_todd 6y agoI would have liked to know what the Kolmogorov approach has to say about the examples used to deflate frequentism ("which of my friends will start a business?"). I don't see from the article how the "smallest-program" approach could say anything useful about those, either. Maybe that wasn't the point--but after poking holes in the frequentist view, it uses unrelated examples like digits of pi to illustrate the Kolmogorov idea, so I'm left unable to directly compare the kinds of statements the two approaches can make. Even going back to dice or coins would have helped me compare them. Like, I know frequentists can show how the variance in coin-toss outcomes decreases as the sample size increases. What can the smallest-program approach say about that? Or was the point that those variant outcomes aren't "real" enough to talk about? Does that mean there is a connection to constructivism in mathematics here? It seems either approach benefits from more data, and there must be a concept related to a "confidence interval" where, as 100, then 200, then 300 digits of pi roll in, your pi-program stays the same size while other programs have to keep growing to accommodate the new data. Like, the ratio of the smallest program to the naive encoding ought to say something about how potentially predictive the small program is. Thanks for the interesting article. It definitely made me think about the issues, and now I'm curious to know more about the topic.
- EbTech 6y agoThanks! I hope to better address your concerns in Part 2.
- PeterisP 6y agoOne aspect of Kolmogorov approach is that implicitly models things like biased probabilities through compressibility. The shortest representation of rolls of fair dice or coins is their exact results, but if there's "less randomness" in some way (biased coin/die, sum of two dice which means non-uniform probabilities, combination of some predictable pattern with random noise) then there are more compact representations of that information, and all of that gets captured by the Kolmogorov approach without any explicit handling of the various possibilities.
- analog31 6y agoMy only thought is that discovering a workable definition of "scientific method" is a reach. Philosophers spent a century searching for such a thing, in vain. On the other hand, providing something that just works would be beneficial enough, even if falling short of the philosophical holy grail, so it's worth pursuing. I'm a physicist, and physicists have always wondered why math works so well in physics. There's this famous essay by Eugene Wigner: https://www.dartmouth.edu/~matc/MathDrama/reading/Wigner.html https://www.dartmouth.edu/~matc/MathDrama/reading/Wigner.htm...
- topsycatt 6y agoThere were a few points that I, as someone unfamiliar with many of the ideas presented, got hung up on. First, the paragraph that begins with "At first blush, the requirement to use..." Seems to be a non sequitur. I don't fully understand how the previous section creates a requirement to use deterministic programs, so I could use more explanation on how that requirement is established. Second, a very simple concrete example of what one of these programs would look like would be immensely helpful. After re-reading the article a bit I have a mental image of a program that contains a long, compressed string and a decompression algorithm that somehow models the system you're interested in. I can imagine how you might get a useful interpretation of probability from the decompression system, but there are enough open questions there that I'm not sure I have the correct interpretation. Hope that helps!
- EbTech 6y agoThanks. I should clarify that the computer is deterministic, so as to avoid building randomness into the definition of randomness! I skimmed over an example too quickly, but your intuition is about right. For that sequence, two possible programs are: - Compute and print the first 40 digits of pi. - Decompress the following string according to a Shannon code with probabilities (1/36,1/18,1/12,[etc]): [insert code]
- carapace 6y agoThis is awesome! Good work! Are you going to touch on Chaitin's Omega?
- EbTech 6y agoThanks! :) I wasn't planning to go there! While I enjoy the idea, for now I'm trying to focus on what's needed to make sense of the problem of induction. Is there a nice connection that I missed?
- danabo 6y agoIt sounds like you are arguing that i.i.d. frequentism doesn't work. I view AIT as generalizing frequentism to non-i.i.d. timeseries. This is formalized as Martin-Lof tests for randomness, and Solomonoff induction.