5 ms·
Here’s a question I had when I learned this stuff in stats class: What does a 95% confidence interval actually mean? I could never get a clear answer.
by rand_r 7y ago
Here’s a question I had when I learned this stuff in stats class: What does a 95% confidence interval actually mean? I could never get a clear answer.
- hombre_fatal 7y agoIt's the age of information. Putting "95% confidence interval" into Google gives back all sorts of examples. What part of it are you having trouble with? Here's a pick at random: https://www.youtube.com/watch?v=mizNAmeAT5w https://www.youtube.com/watch?v=mizNAmeAT5w ("What Does a 95% Confidence Interval Mean?")
- bonoboTP 7y agoNot sure if the answer will be clear, but that's inherent in the weird and counter-intuitive nature of the concept itself (which I consider as a major disadvantage of it). It is an interval that was obtained using a confidence interval generation procedure, for which the following is true: for any true value of the quantity you're estimating, the procedure yields an interval that contains that correct value 95% of the times, when run on different random samples collected conditioned on that true value. You want to estimate some quantity, let's say the cancer risk ratio of hair-dyers to non-hair dyers. You make some (noisy) experimental measurements on people. Then, let's say someone hands you a confidence interval generating procedure. You feed it with your measurement data and it spits out an interval. To claim that this interval is a 95% confidence interval is actually a claim about the confidence interval generating procedure, not about the particular interval. And that claim is that for any true cancer risk ratio value, if we ran lots of experiments and fed all the results to this procedure, then 95% of the resulting intervals would contain the correct corresponding value. So in more detail: Assume the true ratio is, say, 1.05. Then if we repeated our noisy experiment in a 1.05-ratio world and applied the confidence interval generating procedure to the results, we would expect to get 95% of those intervals containing the number 1.05 and 5% of them not containing it. If we assume the ratio is 1.2, the same thing should hold. For this analysis to work, we need an idealized measurement model that tells us what the distribution of measurements looks like given any particular underlying value of the quantity of interest (say the cancer ratio). Then using this model and the description of a given confidence interval generating procedure, we may be able to prove mathematically that it indeed has the above property. Most of the time, statisticians use standards procedures, for which this property has been proved and the proof is widely known. They rarely invent new such procedures. So, in summary: we don't know what the true value is in reality. But we can prove that the confidence interval generating procedure has the above explained property. Then, we will call the interval that this procedure gives us on our real measurement data a 95% confidence interval. What it definitely does not mean is that the true value is in the interval with 95% probability, and that's the most common misunderstanding. I'm not sure if this helps at all. These things are really counter-intuitive and twisted against our natural way of thinking.
- rand_r 7y agoFraming it as statement about an interval generating procedure is really interesting! Do you know what an example would be? Appreciate anything you can point me towards to learn more. Honestly, I feel like the whole concept is straight up weird and fascinating.
- akvadrako 7y agoIt means 95% of the studies should turn out to be a real effect, not just a random coincidence.
- throwawayhhakdl 7y agoSimplest terms: Our sample comes from a true population that we don’t know. Looking at the data we got (primarily the variance and the sample size) we try to infer what the true population mean is. If the spread is small and the sample size is large, odds are pretty good that our sample mean is close to the population mean. If the sample size is small and the variance in the sample is really high, then odds are good our sample mean is pretty far from the true population mean. We can quantify that. Basically saying, if the true population mean is x, how improbable would a sample as different as ours have been? When we construct a 95% confidence interval we’re saying “assuming the randomness of my sample is no more improbable than a 1/20 dice roll, the true population should be this far or closer to my sample mean.” For an intuitive example, consider a 100 sided die going from n to n+100 (e.g. 7 to 107) where we are trying to guess the middle value (n+50) based on a single roll. We get 100. If we want to be 100% confident, the then confidence interval should be 0 to 200 as that’s the range the true value must be in. If we say, nah I’m sure I got it exactly right, your 0% confidence interval is just 100 to 100. If you put on your scientist hat, you say you doubt your particular roll was in the top or bottom 5% of the die, so let’s call it a 90% confidence interval at 5 to 195. Note that you don’t have any particular reason for that last statement, and you could be wrong. In other words a 95% interval says here’s the range that the true population mean should be in assuming our particular sample was no more improbable from the truth than 1 in 20. If the null result is in those bounds, in the article, 1.00 (for no effect) then even though your sample mean might be 7%, your data says it wouldn’t be outrageous to get a sample mean of 7% by chance from a true population mean of no effect.
- cameldrv 7y agoHere's the way I understand it: 1. The model says there exists some true value of a parameter, say, how much more likely you are to get cancer if you use hair dye. 2. You run an experiment with two groups, hair dye, and no hair dye. 3. If you were to run this experiment 100 times, you would expect that 5 of those times, the true value wouldn't be within the confidence interval, and 95 times, it would be. This type of analysis doesn't tell what you really want to know: how likely is it that hair dye causes cancer. What you do know is that you should expect that if you run a bunch of experiments, you should expect 5% of your 95% confidence intervals to be "wrong." To have a degree of belief in something, you need a prior. If I run a hazard ratio analysis for skydiving without a parachute, and the 95% CI includes a hazard ratio of 1.0, I'm still not going to believe that this is a safe activity. If I run it on drinking a glass of water, and the 95% CI doesn't include 1.0, I'm still going to believe that drinking a glass of water is safe.
- otabdeveloper4 7y agoHeh. It means that there's a 5% probability that a result is outside the interval purely by random chance.