6 ms·
I'm too lazy to do the math, but I'm just going to say a test like this needs more than a sample size of 12 to show significance for detecting something that ha
by codeonfire 11y ago
I'm too lazy to do the math, but I'm just going to say a test like this needs more than a sample size of 12 to show significance for detecting something that happens 1 in 500. The woman also was predisposed to thinking that at least some of the 12 had Parkinson's and the sample selection ensured that at least half did have Parkinson's. An actual test would need to allow any possible sample including those that had zero Parkinson's patients.
- comex 11y agoWell, to be conservative, assume she was told or guessed that 7 out of the 12 samples were from patients with Parkinson's, which more than accounts for the "predisposed to thinking at least some" factor - actually it would make more sense to guess six, but that can only make the correct response less likely. Twelve choose seven is 792, so out of the 792 possible responses, she chose the single exactly correct one. No fancy statistics needed - the p-value (defined, according to Wikipedia, as the probability of obtaining an equal or more extreme result, the latter of which doesn't apply as the result is already the most extreme possible for the experiment) is simply 1/792 or about 0.001 (0.1%).
- codeonfire 11y agoThat's ridiculous. That's like saying someone who wins the lottery is clairvoyant with P=0.00. I mean they got the exact right numbers which had 1 in 147 million chance, right?
- exgrv 11y agoA lot of people play the lottery. Only one person did this experiment. When you test multiple hypothesis (as in your lottery example), you need to perform a correction[2,3]. [1] https://xkcd.com/882/ https://xkcd.com/882/ [2] https://en.wikipedia.org/wiki/Multiple_comparisons_problem https://en.wikipedia.org/wiki/Multiple_comparisons_problem [3] https://en.wikipedia.org/wiki/Bonferroni_correction https://en.wikipedia.org/wiki/Bonferroni_correction
- swsieber 11y agoThat's totally different. They didn't test a whole bunch of people a whole bunch of times. That's what the lottery is. That's different math.
- codeonfire 11y agoHow do I know that? I just assume the guy that won was the only guy that played. If we can forget about all the people who ever claimed to smell disease, then we can forget about all the other lotto players.
- oldmanjay 11y agoWhat I find amusing here is the only intellectually honest way you have out of the logic hole you've dug is to claim that you understand neither the lottery, nor the article, which then makes your opinion on either uninformed and not terribly compelling. Were you just advocating for the devil?
- klodolph 11y agoYou're right... you are too lazy. Traditional null hypothesis would be something like "she guesses right 50% of the time", which gives a likelihood of 1/16384 that she would get the correct answer. Let's be cynical, and suppose that she knew or guessed that there were 5-7 patients with Parkinson's, the likelihood is now 1/2538, still pretty low. Even if she knew there were exactly 7 patients with Parkinson's (quite a cynical null hypothesis!) the likelihood is only 1/792. Hey, that's a p-value of 0.0012! Yes, p-values suck. But, the significance is absolutely there, but we would want to follow this research up with a larger sample size and control more of the variables.
- sliverstorm 11y agoI do love how counterintuitive sample sizes are. I swear I see the "it's intuitively obvious that 45 people is a comically small sample size, the researcher is an idiot and this is useless!" comments all the time. Then somebody does the math and it turns out p ≈ 0
- klodolph 11y agoYeah... and then there's the people who benchmark computer programs, don't even bother measuring the test variance, don't know about warming the cache, and post raw timing data online claiming that X is better than Y. Dunning Kruger is a harsh mistress.
- Gravityloss 11y agoLet's apply Bayesian reasoning. :P P(skill): Let's choose a prior that someone can smell parkinson's as one in a million, or 10^-6. (Not very well argued, I admit.) P(data): This exact data's random occurrence probability is 1/2538. P(data|skill): The probability of the result, taking account that she has the skill, is 1 (this assumes she never errs). So we get P(skill|data) = P(data|skill) x P(skill) / P(data) P(skill) = 1 x 10^-6 / (1/2538) = 0.002538 Or 0.25 percent probability, based on this test, that she has the skill. Which is low. Intuitively, I would have expected the calculation to yield a much higher number. The prior was very low though. So I think the grandparent post has some merit. It can be argued that the claim is so extraordinary (the prior) that even twelve "coin tosses" guessed right in a row is more likely.
- molyss 11y agoI believe there are some treatments out there aimed at slowing down the progress of the disease. Diagnosing early might mean slowing down the nasty stuff earlier, thus improving the life of the patient. Also, this sounds like something that is not at all part of the known symptoms of disease. Imagine if the change of smell is linked to something that is a cause for parkinson (I'm not suggesting the article even considers that option, but it's not impossible), and imagine that that something is actually "easy" to treat for. That would change a lot of things ! Yes, it's very unlikely, but even a small increase in the knowledge of a disease is always progress.
- phaemon 11y agoThe odds of randomly guessing all 12 correctly is 1 in 2^12, or 1 in 4096.
- codeonfire 11y agoThe odds of randomly drawing a royal flush are 1 in 649,000 but it happens all the time.
- phaemon 11y agoNo, it doesn't happen all the time. It happens, on average, 1 time in 649,000. That's what those odds mean.
- adenadel 11y agoThat's because there are hundreds of thousands or millions of draws occurring. You need to account for what's called the multiple testing problem. [1] Some ways to do this are to control the family-wise error rate (FWER) or the false discovery rate (FDR). You can control the FWER with the Bonferroni correction or the Sidak correction and you can control the FDR with the Benjamini-Hochberg procedure. You are completely correct about things like this happening by chance, and this is the cause of publication bias in science (since there is a predisposition for publishing positive results). 1. https://en.wikipedia.org/wiki/Multiple_comparisons_problem https://en.wikipedia.org/wiki/Multiple_comparisons_problem
- adenadel 11y agoThis assumes that people 50% of people have Parkinson's and that the people are chosen randomly from the population. Unfortunately neither of those things are true.
- codeonfire 11y agoHere is a table relating the sample size, sensitivity, and confidence interval for evaluating diagnostic medical tests http://www.nature.com/nrmicro/journal/v8/n12_supp/fig_tab/nrmicro1523_T2.html http://www.nature.com/nrmicro/journal/v8/n12_supp/fig_tab/nr... As you can see, a sensitivity of 95% (true positive 95% of the time) with a confidence interval +- 4.3% requires 100 positive subjects. Because the natural rate is 1 in 500, we would need 500*100 = 50,000 total subjects. So you can see how absolutely ludicrous it is to say a woman sniffs the clothes of 12 subjects and is presumed to have a 100% true positive rate with 100% confidence level.
- KingMob 11y agoYou clearly need to take stats again, because you don't understand what you're citing. A 95% confidence interval is not the same thing as a hypothesis test of p<.05. A confidence interval is for estimating the uncertainty that your chosen margin of error will include the true population parameter. Nobody's claiming she's always 100% accurate. That's also a factor of low sample sizes. But I went ahead and computed the margin of error for you. For n=12, a 95% confidence interval requires a margin of error of 28%, so her true detection ability is, at worst, 72%, which is still higher than anything else we've got.