4 ms·
In fact, it is accounted for. You'll notice the negative exponent on the unlikely hypotheses is getting pretty extreme after a handful of flips.
by dthunt 13y ago
In fact, it is accounted for. You'll notice the negative exponent on the unlikely hypotheses is getting pretty extreme after a handful of flips.
- bermanoid 13y agoIn case it's not clear (that wording doesn't really click for me, since there are no exponents involved), the way this information is encoded is in the shape of the distribution itself. If there have been very few observations, it will be very wide, but after many it will narrow to a tiny spike. In this context, at least, the prior distribution encodes everything - there's no meaning to the idea that you're more or less confident in the prior, because the prior already represents your uncertainty about the outcomes. If you were 50/50 on this prior versus another one, then your actual prior would be the average of the two. There is a subfield of statistics that deals with imprecise probabilities, but that's a whole other can of worms and doesn't really relate to this problem. That said, it's fascinating stuff, and very useful in some contexts (if you're uncertain about your priors, it can be useful to do sensitivity analyses to figure out exactly how the end result depends on your prior).
- dthunt 13y agoWhat I meant was that for h_1%, you're going to wind up with probabilities in the area of like, 1.2e-34, very quickly, if the coin is a fair one. The update process is adjusting via multiplication, and for a really bad hypothesis, that's going to bring its probability dramatically close to zero without that many trials. Even though the absolute difference between 1e-4 and 1-e34 hypothesis feels smallish when you look on a linear scale from 0-1, a 1-e34 is a lot 'stickier'. Your explanation has the benefit of being a better explanation; my goal was just to explain where the inertia was hiding.