8 ms·
Beautiful Probability
- deleted 3y ago[deleted]
- d0mine 3y agoBayesian approach sounds like a religion (one true way). There is nothing unusual about different mathematical methods/models producing different results e.g., the number of roots even for the same quadratic equation may depend on "private" thoughts such as whether complex roots are of interest (sometimes they do/sometimes they don't). All models are wrong some are useful.
- usgroup 3y agoYeah I’d agree at some depth. We don’t talk enough about integers, rationals and real numbers and what they imply for our “normative rationality” or “epistemological commitment”. But aside from the integers, everything else is totally suspicious.
- biomcgary 3y agoOne of my priors: "a group of people who look like a faith-based community, but claim not to be one, should not be trusted".
- lalaithion 3y ago> the number of roots even for the same quadratic equation may depend on "private" thoughts such as whether complex roots are of interest You are confusing ambiguity in a problem statement due to human language being imprecise with two well-specified identical experimental results having different results due to the intentions of the human carrying them out. Is arithmetic a religion because there's "one true way" of adding integers?
- kevindamm 3y agoI can think of at least two ways to add integers.. the categorical way that applies a mapping from the set into itself, and the set-theoretic way that deals with unwrapping and rewrapping successor relations. The latter is sometimes resorted to in heavily-relational contexts like Datalog.
- lalaithion 3y agoYes, this is addressed in the original article... there are multiple "lawful" ways of adding integers which all give the same results, and likewise in probability all "lawful" ways of analyzing data should give the same results. If you have two different ways of adding numbers which give different results, one is not lawful.
- d0mine 3y agoIt is not about human language being imprecise. I can formulate the questions using the math language precisely with the exact same result (different number of roots are possible for different formulations of the problem for the same "physical" (coefficients of the quadratic equation) setup). The Map is not the Territory. Different maps can be useful. No true map.
- pdonis 3y ago> the number of roots even for the same quadratic equation may depend on "private" thoughts such as whether complex roots are of interest No. The number of roots that you care about might depend on your private thoughts; but the number of roots itself does not. It's a mathematical fact. It just might not be a mathematical fact that you actually care about. But what you care about is not part of math.
- lupire 3y agoWhy are you ignoring the quaternion roots? 3x3 matrix roots?
- pdonis 3y agoNormally "quadratic equation" means "over the complex numbers" (and the mention of "complex roots" in the post I responded to bears out that interpretation). But yes, different mathematical models can give different answers for things like "number of roots of an equation". But that doesn't mean math depends on "private thoughts". It just means you need to specify which mathematical model you are talking about.
- pdonis 3y ago> Bayesian approach sounds like a religion (one true way). Only about the things that can be mathematically proven. Which is just like any other branch of math. It is true that some Bayesians (and EY can be argued to be among them) like to talk as though Bayesian computation is a drop-in replacement for your brain. Of course it isn't, and Bayesianism, like any mathematical approach, should be taken with a good-sized dose of humility. As Bertrand Russell said, to the extent that mathematical propositions refer to reality, they are not certain, and to the extent that they are certain, they do not refer to reality.
- jawarner 3y agoIsn't that Edwin T. Jaynes example just p-hacking? If only 1 out of 100 experiments produces a statistically significant result, and you only report the one, I would intuitively consider that evidence to be worth less. Can someone more versed in Bayesian statistics better explain the example?
- skulk 3y agoI find the original discussion to be far more interesting than whatever I just read in TFA: https://books.google.com.mx/books?id=sLz0CAAAQBAJ&pg=PA13&lpg=PA13#v=onepage&q&f=false https://books.google.com.mx/books?id=sLz0CAAAQBAJ&pg=PA13&lp...
- usgroup 3y agoYeah generally Jaynes book is very nice and easy to read for this sort of material.
- abeppu 3y ago> One who thinks that the important question is: "Which quantities are random?" is then in this situation. For the first researcher, n was a fixed constant, r was a random variable with a certain sampling distribution. For the second researcher, r/n was a fixed constant (approximately), and n was the random variable, with a very different sampling distribution. Orthodox practice will then analyze the two experiments in different ways, and will in general draw different conclusions about the efficacy of the treatment from them. But so then the data _are_ different between the two experiments, because they were observing different random variables -- so why is it concerning if they arrive at different conclusions? In fact, the _fact that the 2nd experiment finished_ is also an observation on its own (e.g. if the treatment was in fact a dangerous poison, perhaps it would have been infeasible for the 2nd researcher to reach their stopping criteria).
- usgroup 3y agoWell no because it’s talking about either a fixed sample size or stopping when a % total is reached. Neither imply a favourable p-value necessarily. I think the author means to say that it’s two methods incidentally equivalent in the data they collect that may draw different conclusions based on their initial assumptions. Question is how do you make coherent sense of it. At level 1 depth it’s insightful. At level 2 depth it’s a straw man. At level 3 depth, just keep drinking until you’re back at level 1 depth.
- usgroup 3y agoSo you know when you believe something and then you update your belief because you get some evidence? Yeah, and then you stack some beliefs on top of that. And then you discover the evidence wasn’t actually true. Remind me again what the normative Bayesian update looks like in that instance. Unfortunately it’s turtles all the way down.
- nerdponx 3y ago> you discover the evidence wasn’t actually true Not really going to vouch for the normative Bayesian approach, but you might just consider this new (strong) evidence for applying an update.
- crdrost 3y agoThe precise claim (I believe) is that the prior update which you had, made some assumptions about the correct way to phrase your perceptions. That is, you say, for the update, "the probability that this trial came out with X successes given everything else that I take for granted, and also that the hypothesis is true" vs. "the probability that this trial came out with X successes given everything else that I take for granted, and also that the hypothesis is false." So you actually say in both cases the fragment, "this trial came out with X successes." What happens if it didn't really? Well, the proper Bayesian approach is to state that you phrased this fragment wrong. You actually needed to qualify "the probability that I saw this trial come out with X successes given ...", and those probabilities might have been different than the trial actually coming out with X successes. OK but what happens if that didn't really, either. Well, the proper Bayesian approach is to state that you phrased the fragment doubly wrong. You actually needed to qualify it as "the probability that I thought I saw this trial come out with X successes given...". So now you are properly guarded, like a good Bayesian, against the possibility that maybe you sneezed while you were reading the experiment results and even though you saw 51, it got scrambled in your head and you thought you saw 15. OK but what happens if that didn't really, either either. You thought that you thought that you saw something, but actually you didn't think you saw anything, because you were in The Matrix or had dementia or any number of other things that mess with our perceptions of ourselves. So you, good Bayesian that you wish to be, needed to qualify this thing extra! The idea is that Bayesianism is one of those "if all you have is a hammer you see everything as a nail" type of things. It's not that you can't see a screw as a really inefficient nail, that is totally one valid perspective on screwness. It's also not that the hammer doesn't have any valid uses. It does, it's very useful, but when you start trying to chase all of human rationality with it, you start to run into some really weird issues. For instance, the proper Bayesian view of intuitions is that they are a form of evidence (because what else would they be), and that they are extremely reliable when they point to lawlike metaphysical statements (otherwise we have trouble with "1 + 1 = 2" and "reality is not self-contradictory" and other metaphysical laws that we take for granted) but correspondingly unreliable when, say, we intuit things other than metaphysical laws, such as the existence of a monster in the closet or a murderer hiding under the bed or that the only explanation for our missing (actually misplaced) laptop is that someone must have stolen it in the middle of the night." You need to do this to build up the "ground truth" that allows you to get to the vanilla epistemology stuff that you then take for granted like "okay we can run experiments to try to figure out stuff about the world, and those experiments say that the monster in the closet isn't actually there."
- AbrahamParangi 3y agoI'm confused in that I don't see how this is troubling. Yes, the two experimenters rolled dice and got the same result, but it's as if one of them was rolling a 6 sided die and the other a 20 sided one. Each experiment is not a result per se but a sample from a distribution. How you infer the shape of that distribution based on the experiment is a function of the distribution of all courses your experiment could have taken. This set of paths is different in each case, which means the inference we make must also be different. There is no inconsistency. The confusion seems to be in assuming that the experimental result was a true statement about the nature of the world rather than a true statement about simply what happened. edit: This seems to me to be a specific case of a general class of difficult thinking where you ask yourself: "what are all the worlds that I might be in that are consistent with what I'm presently observing".
- lalaithion 3y agoIf you see two people roll a d20 and get a 20, you get to say "wow, that was unlikely" to both of them, even if one of them privately admits they were going to quickly re-roll their die if they got below a 10. What matters is their actual behavior (identical in the example) not their intentions. The d6 vs d20 version is different because their behavior is different.
- ninthcat 3y agoUnlikely in what probability space? We only see one version of reality so the probabilities that we assign to any outcome are based on a prior choice of probability space. That is why the researchers' intent matters.
- AbrahamParangi 3y agoYes, indeed.
- lalaithion 3y agoBoth events have the same probability of happening; 1/20. The fact that the researcher intended to do something in a reality that didn't happen isn't relevabnt.
- birdofhermes 3y agoAs other commenters have pointed out any given introductory chapter in a book on Bayesian statistics, including Jaynes’, is better exposition than this. I found _Probability Theory: The Logic of Science_ very easy to follow and very well-written. I had a similar experience when I finally found a copy of Barbour’s _The End of Time_ and discovered, much to my chagrin, that it wasn’t nearly as mystical or complicated as EY makes it seem in the Timeless Physics “sequence”. Barbour’s account was much more readable and much easier to understand. Yudkowsky just isn’t that great of a popular science writer. It’s not his specialty, so this shouldn’t be surprising.
- lalaithion 3y agoHere's a link: http://www.med.mcgill.ca/epidemiology/hanley/bios601/GaussianModel/JaynesProbabilityTheory.pdf http://www.med.mcgill.ca/epidemiology/hanley/bios601/Gaussia... And if you want to read what he has to say on the optional stopping problem, you can scroll down to page 196 (166 in page numbers) to the heading "6.9.1 Digression on optional stopping" I don't personally think Jaynes is much easier to read than Yudkowsky, but he's definitely more rigorous.
- xelxebar 3y agoJaynes is great, but The Logic of Science is a bit rough around the edges, with lots of errata. Jaynes died when the book was really just a very rough draft plus notes. Bretthorst had to go in and turn it into something publishable, not an enviable task by any means. Here's a list of errata and commentary, collected by a fan: https://ksvanhorn.com/bayes/jaynes/index.html https://ksvanhorn.com/bayes/jaynes/index.html.
- topologie 3y agoThank you for this! I had spotted some errors here and there, but it's always good to have them in one place. I think we are all in the same wagon when I say that even with those rough edges Jaynes' book is kind of a transformative experience for everyone who has already been "conditioned" to other Probability texts. For example, for me Feller is a great intro to "start working with Probability," but Jaynes is where one starts actually "thinking in Probability." The whole Maximum Entropy thing was mind blowing for me.
- bdjsiqoocwk 3y agoMeaningless drivel.
- lalaithion 3y agoFrom _Probability Theory: The Logic of Science_: > Then the possibility seems open that, for different priors, different functions r(x1,..., xn) of the data may take on the role of sufficient statistics. This means that use of a particular prior may make certain particular aspects of the data irrelevant. Then a different prior may make different aspects of the data irrelevant. One who is not prepared for this may think that a contradiction or paradox has been found. I think this explains one of the confusions many commenters have; for an experimenter who repeats observations until they reach their desired ratio r/(n-r), the ratio r/(n-r) is not a sufficient statistic! But when we have an experimenter who has a pre-registered n, then ratio r/(n-r) is a sufficient statistic. However, in either case, > We did not include n in the conditioning statements in p(D|θ I) because, in the problem as defined, it is from the data D that we learn both n and r. But nothing prevents us from considering a different problem in which we decide in advance how many trials we shall make; then it is proper to add n to the prior information and write the sampling probability as p(D|nθ I). Or, we might decide in advance to continue the Bernoulli trials until we have achieved a certain number r of successes, or a certain log-odds u = log[r/(n − r)]; then it would be proper to write the sampling probability as p(D|rθ I) or p(D|uθ I), and so on. Does this matter for our conclusions about θ? > In deductive logic (Boolean algebra) it is a triviality that AA = A; if you say: ‘A is true’ twice, this is logically no different from saying it once. This property is retained in probability theory as logic, since it was one of our basic desiderata that, in the context of a given problem, propositions with the same truth value are always assigned the same probability. In practice this means that there is no need to ensure that the different pieces of information given to the robot are independent; our formalism has automatically the property that redundant information is not counted twice.
- roenxi 3y agoThat seems a bit long winded since this situation is a direct result of Bayes' theorem. It seems to me equivalent to say: Bayes' Theorem holds because it can be proven. Therefore, situations can be constructed where considering identical data without considering priors gives nonsense conclusions. For example if we happen to know as a prior that P(outcome of experiment is a certain ratio) = P(experiment is completed) then that must be considered when interpreting the results.
- 4bpp 3y agoI think there is a simple solution to the thought experiment in the beginning, ignoring the paragraphs upon paragraphs of EY liking the sound of his own voice: The information content of each experiment consists of more than just the stated number of patients tested and success rate. In particular, each experiment report I notice is strong evidence that someone actually used humanity's limited resources to perform that experiment, and slightly less strong evidence that they actually followed the stated procedure. Therefore, the completion of the "stop when I have a high enough success rate" experiment should cause me to update in favour of people with the means to actually running such an experiment, and hence make it more likely that at this very moment there are other research groups out there that are like 1000 patients in and have not yet gotten their 60% success rate.
- mturmon 3y agoThis essay is so weird to read. The author is extremely passionate, yet also claiming to be simply rational. He’s throwing terminology around (Dutch book, ZF) but seems unaware of the limits of the approach he advocates. There are so many cracks in the Bayesian edifice promoted in TFA! These problems are well-known in the Theories of Probability community [1] (which is only a subset of the larger set of theorists recognizing the limits of mechanical Bayesian reasoning in decision problems). Here are a couple. (1) Bayesian approaches force you to assign a sharp probability to every event. How do we map any event to a sharp probability? E.g., I need to give a number for the probability of rain tomorrow, a non-repeating event. How do I map that to a number? Not through relative frequencies- it’s non-repeating. If two people give different numbers, how do we decide who is right? This problem is what Peter Walley has called the “Bayesian dogma of precision.” [2] (2) As noted above in an aside, we have a hard time computing probabilities. This is a practical problem that we all are aware of, but often discount. In what we could call CMP (Conventional Mathematical Probability - Kolmogorov’s axioms) we typically can’t even correctly enumerate the sample space. We’re always forgetting something, so our models are too confident. (In the “Dutch book” analogy alluded to in TFA, we are following the axioms but are somehow always losing money, in a very real sense.) Related to this problem of computing probabilities, we don’t have a rigorous way to determine when two real-world events are independent. Yet we constantly invoke independence to construct models. Kolmororov’s 1933 manuscript was clear on this problem. [3] Not satisfied with this, we go on to hypothesize conditional independence relationships in order to feed our complex “rational” Bayesian machine. It’s thirsty for numbers, and we just make them up! * This all sounds somewhat hypothetical. It’s not. In my day job, I compute supposed Bayesian credible intervals for various physical variables. The people downstream who use those variables to assimilate into physical models typically multiply our credible intervals by 2. My friend across lab has it even worse, they multiply his Bayesian intervals by 3. This is not a well-functioning machine. [1] E.g., https://isipta23.sipta.org/ https://isipta23.sipta.org/, or https://plato.stanford.edu/entries/imprecise-probabilities/#NonCha https://plato.stanford.edu/entries/imprecise-probabilities/#... [2] https://issuu.com/impreciseprobabilities/docs/imprecise_probabilities https://issuu.com/impreciseprobabilities/docs/imprecise_prob..., first paragraph, although the whole short article is on-point [3] from memory, the quote is something like, “determining the conditions under which events may be judged independent is one of the major outstanding problems in theory of probability“
- psychoslave 3y ago>Think laws, not tools But laws are tools, and the esthetical intellectual elegance is an epiphenomenal bonus or a mean to keep human psychism motivated to keep its focus away from all the other attention sinks that life throw at it. And that apply for both law in judiciary and sciences parlances.
- randomsolutions 3y agoI use Bayesian methods often, but this a just religious. Bayesian methods are just that, tools, methods for approaching a problem. There are no laws for applying probability to the real world. To think so puts too much faith in your models. Remember, all models are wrong. Applying probability to the real world requires a host of assumptions, regardless of the methods you use. Frequentist and Bayesian methods have different goals, both have there place. For a counterweight to the strong likelihood principle find discussions of Larry Wasserman: https://youtu.be/Z-YvWyM6dRQ?si=qwzRiaPbj9ruiUEv https://youtu.be/Z-YvWyM6dRQ?si=qwzRiaPbj9ruiUEv And for a balanced discussion for why both are great see Michael Jordan: https://youtu.be/HUAE26lNDuE?si=cwg6wpRS1gXL6r1Y https://youtu.be/HUAE26lNDuE?si=cwg6wpRS1gXL6r1Y