24 ms·
Estimating the chances of something that hasn’t happened yet
- cperciva 8y agoThis reminds me of a different "rule of 3": If you want to compare two things (e.g., "is my new code faster than my old code"), a very simple approach is to measure each three times. If the all three measurements of X are smaller than all three measurements of Y, you have X < Y with 95% confidence. This works because the probability of the ordering XXXYYY happening by random chance is 1/(6 choose 3) = 1/20 = 5%. It's quite a weak approach -- you can get more sensitivity if you know something about the measurements (e.g., that errors are normally distributed) -- but for a quick-and-dirty verification of "this should be a big win" I find that it's very convenient.
- comboy 8y agoI agree that based on combination (n! / k!(n-k)!) it seems to be 1 in 20, but when you think of it as running A and B and checking if X or Y is faster 3 times, then you get 1 in 8. Quite a big difference. Where does it come from? Am I doing something wrong? I mean combination approach counts same time results as a win but given enough time precision we could skip this scenario. edit: Ah, got it, in my case 3rd result for X can be higher than the first result for Y. You said all times must be smaller.
- JoshuaDavid 8y agoLet's say there are 6 trials -- name them x1, x2, x3, y1, y2, y3. In the case of checking if x is faster than y 3 times, you're doing x1 < y1 AND x2 < y2 AND x3 < y3. In the case of checking if all three measurements of x are smaller than all three measurements of y, you're checking x1 < y1 AND x1 < y2 AND x1 < y3 AND x2 < y1 AND x2 < y2 AND x2 < y3 AND x3 < y1 AND x3 < y2 AND x3 < y3. In other words, the latter case is checking whether the slowest x of three trials is faster than the fastest y of three trials.
- comboy 8y agoYou're right. I actually edited before you replied. Should have deleted it or preferably think more before typing. But now it's forever.
- Waterluvian 8y agoI'm glad you kept it. The explanation was enlightening to me.
- rivp931 8y agoGoing down this rabbit hole eventually leads you to nonparametric statistical tests, e.g. Mann-Whitney-U and so on.
- cperciva 8y agoRight. And those are also more powerful (and don't need any knowledge of the distribution of measurement errors). I mentioned the "three old and three new" test because it's simple, not because it's powerful.
- blt 8y agomore powerful in the statistical sense? I though non-parametric tests were usually less powerful than those where a certain distribution is assumed.
- adamc 8y agoThat depends on whether the assumed error distribution is accurate. Obviously, if the distribution is known, you will do better by incorporating it. But assuming Gaussian errors when the data doesn't match can lead to bad analyses.
- srean 8y agoYou would be surprised what number of samples does to the tests. Take t-test -- the most powerful test to check if means of 2 equivariant Gaussians differ. If you compare the asymptotic efficiency of Mann-Whiney (a distribution free test) relative to t-test is around 0.96. Of course in practice you will not have infinite samples. It then comes down when do these asymptotics kick in. Unfortunately that depends on the distribution.
- yen223 8y agoMy personal favourite quick-and-dirty trick: A quick way to estimate any distribution's median is to draw 5 random samples. There's a >90% chance that the median is between the biggest and the smallest value.
- zawerf 8y agoIs this a good method to use while trying out a lot of ideas? Usually when I am optimizing code, most of the ideas don't work out and performance remains roughly the same (or so I think - I don't really know and want a better workflow here). But if you do this test repeatedly, even if the code had identical performance, you'll get a false positive 5% of the time. And depending on the spread of the timings you might not get a clear XXXYYY win even when it is indeed a minor improvement.
- ma2rten 8y agoIt is not a good method to use in that case.
- eloff 8y agosqlite famously squeezed out a ~40% performance improvement (I think from v3 to v4?) by just combining tons of micro-optimizations of this kind where it wasn't obvious if each change even made an improvement. They measured the performance with cachegrind in order to identify very small improvements that get lost in the normal measurement noise.
- zawerf 8y agoI would love to read more about this. Do you have a link?
- SQLite 8y agohttps://www.sqlite.org/cpu.html https://www.sqlite.org/cpu.html
- rytill 8y agoThat’s a really awesome idea. Also, maybe a few of the performance boosts in the final combination are mutualistic, but if measured independently would not make a difference.
- thomasahle 8y agoAlso check fishcooking: http://tests.stockfishchess.org/tests http://tests.stockfishchess.org/tests which is how Stockfish became the indisputably strongest chess engine: Everyone can commit patches which are then tested in thousands of games, and if they make Stockfish stronger, they are accepted.
- luckyt 8y agoThis is incorrect: there's no reason to expect that X and Y will each appear 3 times in 6 trials if their probabilities are equal. If all 3 measurements of X are smaller than all 3 measurements of Y, then you have X < Y with confidence 1 - 1/8 or 87.5% confidence. You'd need at least 5 measurements to be 95% confident.
- saagarjha 8y agoWell, you’re the one measuring each three times. So this should always be the case?
- bo1024 8y agoYou're not considering the right probability space. We have 3 measurements of X and 3 of Y. The question is the distribution on orderings of these six measurements. If X and Y come from the same distribution then all orderings are equally likely.
- gerdesj 8y agoI'm also having trouble with this. On the face of it the quick and dirty "XXXYYY" test outlined above looks good but are these two following statements consistent? ie are the run times of X (new code) and Y (old code) really from the same distribution. "is my new code faster than my old code" "If X and Y come from the same distribution then all orderings are equally likely"
- bo1024 8y agoI'm thinking of it as a statistical hypothesis test. The null hypothesis is that they come from the same distribution. Under that hypothesis, there's only a 0.05 chance of seeing three X tests all below three Y tests. So if we see this, we can probably reject the null. If we think X and Y distributions are both something like normal with similar variance, then we should also be able to say the chance of XXXYYY given Y is better than X is at most 0.05. But if the distributions for X and Y can be really different, then I think you're right -- this test could be misleading! For example, say Y always takes 2 seconds, and X takes 1 second 90% of the time, but 1% of the time it takes an hour. If we run three tests of each, we'll probably only see good runs from X and conclude it's better, when it's not.
- taneq 8y ago> If the all three measurements of X are smaller than all three measurements of Y, you have X < Y with 95% confidence. On a multitasking system, you should use the fastest benchmark run (assuming you're running the same code on the same data in all cases, and if you're not then you're not really benchmarking). This will be the one that is least influenced by any other processing going on.
- gugagore 8y agoThat's the case when you think the measurement noise is only positive. But are there kinds of noise that reduce your benchmark time? On one hand, I believe a piece of code has a true benchmark time, but we only measure noisy versions that are skewed positive due to multitasking noise. On the other hand, it seems important to make benchmark measurements in the context of the whole system. If A beats B in the best case (unloaded system), but doesn't win in the average case (taking the context into account), then that's important information.
- skybrian 8y agoThe way I think about it is that benchmarks include deterministic operations (that happen every time) and nondeterministic operations (that don't). And all of these numbers are positive, since there are few things in the world that take negative time. By taking the minimum, we focus attention on the deterministic operations. This isn't likely to be realistic for production, but production won't look much like your benchmark hardware anyway. When making code changes, deterministic operations are the part you have most influence over and have the most impact. Removing an unnecessary, deterministic operation will speed up every run, not just some of them.
- ced 8y agoMan, I wish I understood frequentist statistics to know if your reasoning makes sense. Bayesianly, if O = "the XXXYYY ordering", and F = "algo B is faster than algo" A, then P(F|O) = P(O|F) * P(F) / P(O) then... what? There isn't even a clear P(O|F) likelihood without making assumptions about the process' noise. If the measurement is very noisy compared to the gain, then XXXYYY is just dumb luck, and doesn't tell you anything. If there is no noise at all, then just getting XY is enough to make a decision.
- justinpombrio 8y agoLet me explain it in a Bayesian way. The Bayesian approach would be to consider a whole bunch of hypotheses about how much better Y does than X (and vice-versa), and a prior distribution over how likely you think each is a-priori. Then for each sample S, you update the probability of each hypothesis H by multiplying its probability by P(S|H), then re-normalize. Well in this case, we're only going to consider two hypotheses. Call them H0 and H1. H0 is the "null hypothesis", and says that X and Y are exactly as fast as each other. H1 is the hypothesis that H0 is wrong and Y is totally faster than X. You'll notice that H0 is oddly specific, and H1 is ill-defined. Don't worry about it. To start off, pick your prior distribution over H0 and H1. Pick whatever you want, because we're going to ignore it shortly. Now some evidence comes in. Time to update! We got the ordering XXXYYY. First, let's update H0. P(XXXYYY|H0) = 1 / 6choose3 = 5%. Wow, that's not a very good update for H0. It's probably just false. For expediency, let's just toss it out. H1 is the remaining hypothesis. H1 wins! Y is faster than X.
- pedrosorio 8y ago“P(XXXYYY|H0) = 1 / 6choose3 = 5%. Wow, that's not a very good update for H0. It's probably just false. For expediency, let's just toss it out.” It only makes sense to toss H0 out if P(XXXYYY|H1) >> 5% (such that the evidence for H1 relative to H0 increases after the observation). You are implicitly assuming that’s the case because “it makes sense”. But as the parent post mentioned, the likelihood is not defined and in particular if the noise in the observation process is large enough, P(XXXYYY|H1) may be very close to 0.05 as well.
- shaki-dora 8y agoI wish the people constantly complaining how n=20,000 is "far too small a sample size to call this science" (for every empirical study) would take not. Effect size matters!
- Fomite 8y agoAn absence of statistical training is dangerous. A little statistical training remains dangerous, but gets annoying as well
- GuB-42 8y agoIn the specific case of code, it is worth noting that runs are typically not independent, because of caching. It is very common for the first run to be slower. It means that by "random chance" slow-slow-slow-fast-fast-fast is more likely than fast-fast-fast-slow-slow-slow. I personally tend to discard first runs as outliers when profiling.
- blowski 8y agoI’m not very good at maths, so I didn’t understand the whole post. However, does the size of the whole population affect the “3/n” thing? For example, if I’ve read 200 pages of a 201 page book and not discovered a typo the chances are 3/200 if I’ve understand the post correctly. If the book has 20000 pages, is the probability still 3/200?
- geetfun 8y agoN is the sample size that you’ve sampled already in your observation. If you’ve sampled all the pages, then we are talking about certainty which this wouldn’t apply.
- mirimir 8y ago> It says that if you’ve tested N cases and haven’t found what you’re looking for, a reasonable estimate is that the probability is less than 3/N. So in your example, according to the rule, the probability of errors is less than 3/200. Which it is, either 0/201 or 1/201.
- mirimir 8y agoThere's a trivial corollary, which I remember from my first physics class. Never base anything on less than three measurements. Or maybe it wasn't even physics, but rather carpentry. The old "Measure twice, cut once." rule is iffy.
- babygoat 8y agoThree data points is not the same thing as three measurements of the same object.
- mirimir 8y agoThere's no question that three measurements of some property of an object are three data points. I vaguely recall that my first physics lab experiment was measuring something with a ruler.
- sonnyblarney 8y ago" Or maybe it wasn't even physics, but rather carpentry. " That's the quote of the day. My father was a carpenter, and every day I think software has more in common with carpentry than computer science.
- Koshkin 8y agoInteresting: the frequentist derivation is using the logarithm, while the Bayesian one, the exponent.
- jey 8y agoBut note that log and exp are inverses of each other, and it's applied to opposite sides of the equation. In particular, this: 1 - exp(-3) ≈ 0.95 Can be rewritten as: -3 ≈ log(1 - 0.95)
- madrox 8y agoWhat the author glosses over somewhat is the method of sampling. If you read the first 20 pages, find no typos, and use this rule to arrive at 15%, that could be way off. He's assuming the risk of typos are evenly distributed when there's a lot of reasons it may not be. For example, the first half of the book could've been more heavily proof-read than the latter half. It's not out of the question that editors get lazier the farther they get into the book. If you were to randomly read 20 pages in a book and find no typos, 15% probability makes more sense. It's understandable to not mention this in a short blog post about the rule of three, but never forget that when you're interpreting statistics...how you build your sample matters.
- bluGill 8y agoIn the case of a book the first pages are likely to be significantly better than the rest. An editor knows that once you have invested in reading the first part of the book you are unlikely to put the book down. This means the first pages have to suck you in, and not do anything to make want to quit reading. The most important part is the first sentence as this is often the only part that drives your buy/leave in the store decision.
- whatshisface 8y ago>The most important part is the first sentence as this is often the only part that drives your buy/leave in the store decision. I usually sample books from the middle. I wonder what the heatmap of bookstore reading really looks like?
- bluGill 8y agoMy informal sampling suggests most people start at first page. They worry that reading the middle or end would reveal a spoiler. It would be interesting to get good data though.
- yummybear 8y agoIn the same vein if the weather is good today, there is s large probability that it’s good tomorrow, but if it’s been good for 14 days there is an increasing probability of bad weather.
- ajkjk 8y agoSo basically the '3' comes entirely from the choice of a 95% confidence interval. If you want a 99% confidence interval it's instead the 'rule of 4.6', which doesn't roll off the tongue as well.
- jedberg 8y agoYou could do the 'rule of 5' though and have a confidence of 99.3%, which is pretty close to 99.
- deleted 8y ago[deleted]
- neolefty 8y agoIf you're truly estimating, though, wouldn't you use a 50% confidence interval — which gives you the rule of 1 (chances of the thing being true are 1/n — if you've seen 20 pages without a typo, chances are 1/20 that there's a typo on a page, with 50% confidence)?
- hawkice 8y agoSo, for a 50% confidence interval, you could look at the first word -- if it is a typo, boom, done. Otherwise, flip a coin. This stinks. Estimating is about approximating an answer with low information -- it's about efficiency of using data, not _only_ doing better than guessing.
- neolefty 8y agoSorry, I meant once you've examined n trials, what is your 50% confidence interval about the odds for a single trial? I think it would be 1/n.
- SamReidHughes 8y agoYes. That's a very reasonable choice.
- deleted 8y ago
- citilife 8y agoI've seen a very similar problem to this referred to as "black swan events"[1]. The whole point is you can't actually compute it. You see the period, it's there. What he's doing here isn't science, it's a guess. As others have pointed it, it's rather rare that events happen with perfectly distributed probability (probably the opposite). For instance, the chance of being in an accident is much higher if an accident just occurred right next to you. It's almost way more likely to get sick, if others are already sick. In fact, although I don't have any statistics to back me up, I'd guess that most events happen in clusters (including spelling errors, or when you test for perfect pitch in children, when you go to the music class). This is essentially a guess, and it's better to say you don't know than guess wildly. [1] https://en.wikipedia.org/wiki/Black_swan_theory https://en.wikipedia.org/wiki/Black_swan_theory
- exp1orer 8y ago> it's better to say you don't know than guess wildly. > although I don't have any statistics to back me up, I'd guess that most events happen in clusters More seriously, yes, before you apply this rule you should think about whether "each event has the same probability of error" is appropriate for your situation.
- vlehto 8y agoI'm nitpicking, but there is not even a chance that this could be "science" with any coherent definitions of science and future. For example there is Karl Poppers falsification standard. Future is never falsifiable as long as it remains in the future. Another approach would be this: Future is impossible to research, as you can only research entities that exist or phenomena that is happening. Future entities and future phenomena do not exist yet ( by definition ) so you cannot research them. -> no science of future. What we are talking here is scientific prediction. "Scientific" is just another word for good and only means something when compared to another method of prediction that can be shown to be worse. So we very much agree. I just wanted to write this out because the Future studies crowd has bothered me for some time now.
- carlmr 8y ago>The whole point is you can't actually compute it. You can't compute the probability. But you can compute an upper bound on the probability with a reasonable confidence, which is what the author is doing here. This is standard statistics and might be useful in some cases. It does assume that the events are independent, but this is a pretty standard assumption that you have to check whether it applies in your case, or at least approximately applies in your case. >What he's doing here isn't science, it's a guess. Which is a pretty big thing in this subfield mathematics called statistics.
- BenoitEssiambre 8y agoThis is cool. The given example of typos on page kind of highlights the fact that more sophisticated math involving a prior might give better results in some cases. The beta(1, N+1) prior is an assumption that you start with the a priori knowledge that a typo rate of 1% and a typo rate of 99% are equally as likely as each other. Most people would assume that books don't have typos on most pages and a 99% typo rate is unlikely. However as your sample gets bigger the prior matters less and less so this rule is still useful. Just know that it is reasonable to bias the results a bit according to your prior when N is small.
- rspeer 8y agoIf the author is reading this: When someone provides you a pro-bono translation, by all means credit them, but do not let them host it on their own site. Frequently, they are siphoning your PageRank. They will eventually replace the translation with monetized content of their choice. A good translation takes work! If you got a translation for free, why should you believe it's good, or that it has no ulterior motive? A firm called "WebHostingGeeks" used to do this all the time. They would offer free translations of blog posts, into languages the author probably didn't speak (and they didn't either, they were just using machine translation). They'd ask authors to link to the translation on their site, and over time they would add their SEO links to the translation. I first noticed this when WebHostingGeeks offered me a Romanian translation of ConceptNet documentation, my roommate spoke Romanian, and he said "maybe I'm not used to reading technical documentation in Romanian but I think this is nonsense".
- gboudrias 8y agoYeah this is weird. Also, if OP doesn't speak Italian themselves, how can they attest to the quality of the translation?
- Recursing 8y agoI'm Italian, for what it's worth the translation seems fine, even if the writing quality is not as good as the source and maybe the joke about "Bayesians being allowed to make such statements" is a bit lost
- thedirt0115 8y agoI have a couple L2’s that I mostly speak and read, virtually no writing - my grammar is poor and my vocabulary is limited. Because of this I would feel uncomfortable/embarrassed to translate any tech stuff I wrote to it. However, I could read someone else’s translation and know if it’s way off the mark. Also, I have plenty of friends that speak the language natively, so I could ask them to review. If they say it’s good, I’d vouch for the translation because I trust them (even if I couldn’t read it at all).
- 8y ago
- motohagiography 8y agoIntuitively, this seems like an obverse of the "optimal stopping problem." https://en.wikipedia.org/wiki/Optimal_stopping https://en.wikipedia.org/wiki/Optimal_stopping or the subset of the Odds Algorithm (https://en.wikipedia.org/wiki/Odds_algorithm https://en.wikipedia.org/wiki/Odds_algorithm) where optimal stopping point to find the highest quality item in a sample is essentially N * 1/e = 0.368... Handwavily, this resembles the rule of three, where we could probably say instead, "the rule of 1/e," because both of these appear to be artifacts of the same relationship and same type of problem.
- the_cat_kittles 8y agoi think you get more bang for your buck if you try to understand the mechanics that generate successes and failures. assuming a flat prior is crazy in almost every real world case. in otherwords, i think effort is probably better spent understanding the problem rather than understanding how to make the most of ignorance.
- TheNewAndy 8y agoThis feels related, and is interesting: https://en.wikipedia.org/wiki/Sunrise_problem https://en.wikipedia.org/wiki/Sunrise_problem (estimating the probability that the sun will rise tomorrow)
- Cyphase 8y agoThis is from 2010.
- simulate 8y agoThis reminds me of an 1999 article in the New Yorker by Tim Ferris called "How to Predict Anything" https://www.newyorker.com/magazine/1999/07/12/how-to-predict-everything https://www.newyorker.com/magazine/1999/07/12/how-to-predict... > Princeton physicist J. Richard Gott III has an all-purpose method for estimating how long things will last. In particular, he has estimated that, with 95% confidence, humans are going to be around at least fifty-one hundred years, but less than 7.8 million years. Gott calls his procedure the Copernican method, a reference to Copernicus' observation that there is nothing special about the place of the earth in the universe. Not being special plays a key role in Gott's method.
- whack 8y agoIn the example given, the author says that the odds of a given page having a typo is less than 3/20. Sure, but if we don't want a range, but an exact number? That sounds like a more interesting challenge to me. Formal statement: - You have observed N events, with 0 occurrences of X - Someone wants to make a bet with you about the likelihood of X happening - Once you've quoted a number, your counter-party then has the option of making an even bet about whether the actual likelihood is greater than or less than your prediction Eg: If you predict 3/N using a 95% confidence interval, then 95% of the time, the actual likelihood will be less than 3/N. Your counterparty will then win the bet 95% of the time, simply by predicting it to be lower. Your ideal strategy would be to quote a likelihood which is over/under 50% of the time, not 95%. Ie, you want to pick E such that 50% of the time, it matches the observation you've made (no occurrences), and 50% of the time it does not. E^N = 0.5 N log E = log 0.5 log E == log 0.5 / N E = 0.5^(1/N) For the example given, that comes out to 0.966. Ie, there's a 96.6% chance of no typos in a given page. Across 20 pages, this comes out to 0.966^20 => 50% chance of no typos. If your goal is to quote the single best estimate which can hold up well in a betting market, I believe this would be the ideal strategy
- olooney 8y agohttps://en.m.wikipedia.org/wiki/Rule_of_succession https://en.m.wikipedia.org/wiki/Rule_of_succession https://en.m.wikipedia.org/wiki/Sunrise_problem https://en.m.wikipedia.org/wiki/Sunrise_problem
- AstralStorm 8y agoThis presumes conditional independency - violations of which are common. Instead, you get to estimate the dependency between each observation as in advanced variants of Bayes chain rule. Ultimately, some place of the estimator will contain an assumption giving only bounded optimality.
- gweinberg 8y agoI don't think you can do this. You have a reasonably good upper bound of the probability, but you don't have any justification for putting any lower bound other than zero on the probability. In particular, if you're considering the probability of a catastrophic event, it's probably better to find some other rationale for estimating the probability than just saying 'it has never happened before'.
- mihaifm 8y agoIf you’re wondering where does the formula (1-p)^n come from, it’s a number often used in gambling (if I throw a die 7 times, what are the chances of getting at least a 3). The probability of an event having probability p happening after n trials is 1-(1-p)^n, and he’s using the inverse of that. https://en.m.wikipedia.org/wiki/Binomial_distribution https://en.m.wikipedia.org/wiki/Binomial_distribution
- davmar 8y ago(1-p)^n is one my favorite things. happy to see it get some attention.
- oldgradstudent 8y agoThe rule of three requires quite a lot of assumptions about the nature of the phenomena. Or as Taleb says it: > Consider a turkey that is fed every day. Every single feeding will firm up the bird's belief that it is the general rule of life to be fed every day by friendly members of the human race "looking out for its best interests," as a politician would say. > On the afternoon of the Wednesday before Thanksgiving, something unexpected will happen to the turkey. It will incur a revision of belief.
- PeterisP 8y agoIt certainly does not require unwarranted assumptions, the proposed approach is consistent with the well-known turkey scenario. From the experience of feeding a turkey can infer that being slaughtered is a rare event, and that the likelihood it happening exactly tomorrow (without having access to a calendar) is not necessarily 0, but is below a certain rate - and it's definitely not likely to happen three times in the next week. Surviving for 180 days is reasonable justification to assume that, on average, turkeys get slaughtered less frequently than every 60 days, i.e. that the likelihood of Thanksgiving suddenly arriving tomorrow isn't 50% but rather something below 2%.
- AstralStorm 8y agoCompletely wrong. Continued survival of given that gives no information at all of survival rates and its distribution. This is because variability of data input is extremely low so mutual information between each of the days is vanishingly small. This is why a good experimental design will observe a measurement with some expected variability. An even better trick is the sleeping beauty problem. To solve it you need external information.
- eximius 8y agoI feel like I'm having a math stroke. The posterior probability of p being less than 3/N for Beta(1, N+1) should be integral(Beta(1, N+1), 0, 3/N), right? That trends toward zero, so I must be wrong, but I can't for the life of me remember why. EDIT: Ah! I was accidentally using Beta instead of the PDF for Beta.
- saagarjha 8y ago> If the sight of math makes you squeamish, you might want to stop reading now. Sigh…another article normalizing the concept that math is something it’s OK to be uncomfortable about…
- YeGoblynQueenne 8y agoSo, according to this rule, if I wait for 10 minutes for a bus to come and none does, and then I wait for another 10 minutes for an alien invasion and none happens, the two have the same upper bound on their probability? Or are we going to start talking about priors, on buses and alien invasions, in which case the rule of three is not really useful? If I want to know how likely a specific book is to have typos, can't I just go look for statistics on typoes in books, and won't that give me a better estimate than a "rule" that will give the same results no matter what it is that it's trying to model?
- PeterisP 8y agoThis is the scenario where you know nothing else other than you waited 10 minutes for this event and it didn't happen. Sure, if you have more data, then you can get much tighter bounds for your estimate.
- nkurz 8y agoThis is the scenario where you know nothing else other than you waited 10 minutes To make the math work, I think you need to make several other assumptions. Don't you also have to know that that ten minutes you sampled are representative? That the events (if they did occur) are independent? It seems odd that you'd in a situation where you can rely on specific assumptions like these while also believing that buses and aliens are equally likely to appear.
- YeGoblynQueenne 8y agoYou always have more data -background knowledge- unless you've only existed in those last 10 minutes. And if something has really never happened before, like my alien invasion example, what have we learned by applying the rule of three? Honestly- perform the experiment yourself. Wait for X time, then calculate 3/X. Do you now have an upper bound on the probability that an alien invasion will happen?
- carlmr 8y ago>Wait for X time, then calculate 3/X. Do you now have an upper bound on the probability that an alien invasion will happen? Assuming we would observe an alien invasion and write about it, we have about 5000 years since humans started writing. So 3/5000a, so an upper bound on an alien invasion in a year (assuming p hasn't changed since we started writing) is 0,06% per year. It's an upper bound in any case, meaning that is this or less. That includes p = 0. You're just 95% confident that it won't be higher than that. But maybe the aliens brought us our writing system.
- brian_herman__ 8y agoMurphy’s law “Anything that can go wrong will go wrong”
- sangd 8y agoThis rule of 3 may be a good example for reading books and finding typos. It is no where as good for estimating the chances of "something that hasn't happened". It's so random using this rule & claiming the result as an estimate.
- throwaway487548 8y agoOh, numeric astrology and probability tantras. Future is not predictable by definition. It is just an abstract concept, a projection of the mind. Any modeling, however close to reality it might seem to be, is disconnected from it, like a movie or a cartoon. Following complex probabilistic inferences based on sophisticated models is like to act in life guided by movies or tantric literature (unless you are Goldman Sachs, of course). For a fully observable, discrete, fully deterministic models, such as dice or a deck of cards probability could only say how likely a certain outcome might be, but not (and never) what exactly the next outcome would be. Estimation of anything about non-fully-observable, partially-deterministic environments is a fucking numeric astrology with cosplay of being a math genius. No matter what kind of math you pile up - equations from thermodynamics, gaussian distributions or what not it is still disconnected from reality stories, like the ones in tantras.
- binarysolo 8y agoMildly interesting trivia: there's a Chinese idiom stating "compare 3 shops for a good deal" (貨比三家不吃虧), I guess that makes sense from a statistical standpoint. :)
- Jedi72 8y agoI was just on a flight, hanging with a guy who's PhD was this topic. Josh is that you?? What are the odds.
- LeonB 8y agoAlthough interesting, the article doesn’t relate to predicting things that haven’t happened yet, just things that aren’t known yet. When predicting things that haven’t happened yet, a publicized “certain” prediction will inevitably influence the actual probability in an unpredictable way.
- jeffdavis 8y agoI'd be interested to know how to estimate things that are very rare or don't have a normal distribution. For instance, let's say I have a bold plan to protect us from meteor strikes that will cost $100B. How would a person make a decision about whether that's a good trade-off or not? And how would a mathematician help them make that decision? How would it change for more complex cases, like a shield to prevent nuclear ICBM warfare which has never happened, but we all are worried about?
- carlmr 8y agoSimple answer, this is a rough upper bound. The more information about your problem you have, the better you will be able to model the probability distribution and the better priors you have on it happening.
- anotheryou 8y agorelated question: probability that the sun will rise again tomorrow https://en.m.wikipedia.org/wiki/Sunrise_problem https://en.m.wikipedia.org/wiki/Sunrise_problem A practical use of a Stilton to this is included in reddits "best" sorting of comments, leaving from room for doubt with low sample sizes of votes. https://redditblog.com/2009/10/15/reddits-new-comment-sorting-system/ https://redditblog.com/2009/10/15/reddits-new-comment-sortin...
- AstralStorm 8y agoUnmentioned assumption of normal distributed errors is pretty evil. This is exactly this approach is worthless for rare events which by definition have very skewed distributions. In that case, probably is overestimated a lot. On the other hand, I'd the is a rare but systematic error, the error probably will likely be grossly underestimated. Thought experiment: suppose you're writing a long string of digits that consists of 1 followed by a large number of zeroes (say 99 for simplicity) followed by (say 10000) uniformly distributed digits. Your writing system has an issue that changes half of 5 digits into 6. What probability of error will be estimated by this dumb method after 100th digit? Correct bayesian approach updates the prior based on input variability keeping the error estimates high when input has low variability etc. (This can be with variance or another method.) The even better method tries to estimate the shape of input distribution. In other words, your result would be a difference of likelihood ratio of both input and output prior distributions (estimated to date - since no errors the ratio would be 1) minus likelihood ratio of posterior distributions.
- rlue 8y agoI don't know the first thing about schools of thought in statistics (frequentist? Bayesian?) but something feels fishy about extrapolating a probability based on sample size (or number of trials) alone. 20 pages, no typos, <15% chance of flawless spelling? What if the manuscript is 2×10⁴¹ pages long? Would someone care to explain why the math says it shouldn't matter (for reasonably small values of p)?
- carlmr 8y ago><15% chance of flawless spelling? per page
- DoctorOetker 8y agoI wouldn't use a probability density of typo's per page, but instead the probability that a word is spelled wrong. Then it's like drawing marbles from a vase containing an unknown proportion of blue and red marbles. I would use the formula (M+1)/(N+2) where N is the number of words, and M is the number of mistakes. Note that for a large corpus (M+1)/(N+2) approaches -> M/N, so we recover the frequentist probability. Also note the author (John D Cook) correctly expresses intuitive doubt that 0 typos / 20 pages can not be construed as certainty of no mistakes. Similarily, seeing a mistake in every word in a subset does not guarantee that all words in the book will have a typo. Let's look at the modified formula (M+1)/(N+2) in these cases: if we observe no typos (M=0) in 1000 words, it estimates the typo probability as (0+1)/(1000+2)=1/1002 != 0, so we can't rule out mistakes. Similarily if all words were typos (M=N) for the first 1000 words we get 1001/1002 != 1, so we can't be sure every future word is a typo. Check out Norman D Megill's paper on "Estimating Bernoulli trial probability from a small sample" where the formula (M+1)/(N+2) is derived, on page 3 it appears: https://arxiv.org/abs/1105.1486 https://arxiv.org/abs/1105.1486
- codeulike 8y agoThis is a useful rule of thumb. Lets try stretching it: Humans havent destroyed the world in their 300,000 years of existence, so probability of them destroying the planet in future is less than 0.00001 (1 in 100,000). I feel like that might be an underestimate. Reminds me of the Doomsday Argument https://en.wikipedia.org/wiki/Doomsday_argument https://en.wikipedia.org/wiki/Doomsday_argument
- fernly 8y agoClearly this is a bad estimating rule for some processes. The one that popped into my mind is the probability of earthquakes (guess where I live...). This would have the probability of an earthquake declining with each quake-free year that passes, where in fact the USGS would say the opposite.
- gugagore 8y agoThere's something that feels different about that. You're talking about the rate of some event occurring (like a Poisson process), and measuring how many occurrences are in an interval (a year). What are the 6 samples you are collecting? 6 years of numbers of earthquakes in a year? Edit: sorry I thought you were replying to another comment.
- johntiger1 8y agoThis post is of course talking about the difference between MLE and MAP estimation. Consider the converse case: you flip a coin once and observe it is heads. Do you then conclude that the probability of heads is 100%? No, because even though you have data supporting that claim, you also have a strong prior belief in what your probability of heads should roughly be. This is encoded in the beta distribution, as mentioned in the article
- Mauricio_ 8y agoThis is very similar to the sunrise problem. Laplace said the probability of the sun rising tomorrow is (k+1)/(k+2), where k is the number of days we know the sun has risen consecutively, if we always saw it rise, and if we don't have any other information. https://en.m.wikipedia.org/wiki/Sunrise_problem https://en.m.wikipedia.org/wiki/Sunrise_problem