4 ms·
Stopping when you hit 95% confidence is a classic failure mode. Yes, if you are doing classic t-test based A/B testing, you have to wait until a pre-determined
by squarecog 10y ago
Stopping when you hit 95% confidence is a classic failure mode. Yes, if you are doing classic t-test based A/B testing, you have to wait until a pre-determined threshold; otherwise, effectively, by looking at the p-value and stopping when it hits 95% confidence, what you are doing is ignoring all negative results and accepting the first positive result you see -- clearly, that's bad science. I'm simplifying the math here, but this is the general notion.
You can see a demonstration of this in practice: http://www.gigamonkeys.com/interruptus/ http://www.gigamonkeys.com/interruptus/ (sit back and watch your false positive rate, aka "bogus a/b testing results", grow).
- tedsanders 10y agoIs it really a failure mode? I think it should be fine to stop whenever you want during a test. You just have to be smart enough to not use a t test and misinterpret it to mean than A is 95% likelier than B.
- daveguy 10y agoYes. It really is a failure mode. No, it is not fine to stop whenever you want during a test. It doesn't matter what statistical test you are using -- you need to gather enough data to have sufficient statistical power. Also, the more tests you do the lower the p-value should be. If you stop when you reach your p-value then you are misinterpreting the results if you claim the results mean anything.
- tedsanders 10y agoThis seems terribly dogmatic to me. Take the example of clinical trials. Suppose you're testing a new cancer drug. You design an experiment to test the new drug, named B, versus an established chemotherapy treatment named A. You expect B's performance to be similar to A's performance in controlling the cancer, so to make sure your trial has high power, you plan to test the drugs on 2,000 patients (with each drug administered to 1,000). Now consider the following two scenarios: (1) After giving drug B to 100 patients, all 100 patients are dead. Do you continue the trial, giving the (apparently) deadly drug B to 900 more patients? (2) After giving drug B to 100 patients, all 100 patients are totally cured (vs A curing 3 in 100). Do you continue the trial, withholding the (apparent) cure for cancer from 900 more patients? In either case, since you have a strong effect, it seems to me there is logical justification to end the trial early. Obviously the stakes are higher in clinical trials than website design, but in both cases, data acquisition has costs and intermediate results may inform changes to your experiment design. I honestly cannot see how anyone could blankly assert that stopping a test is always wrong. There are certainly circumstances where you do want to stop early. You just have to make sure you aren't misinterpreting a statistic when you do so.
- apathy 10y ago"Wrong" isn't the word we're looking for here, I don't think. But your above example is bullshit -- nobody puts 1000 patients at risk in a phase I (safety) trial, and if the dose isn't reasonably well calibrated by the phase III study you're describing above, someone's going to jail. In Phase II we will often have stopping rules for exactly this reason, just in case the sampling was biased in the small Phase I sample. Above there are a number of things to notice: 1) The phasing approximates Thompson sampling to a degree, in that large late-phase trials MUST follow smaller early phase trials. Nobody is going to waste patients on SuperMab (look it up). 2) The endpoints are hard, fast, and pre-specified: IFF we have N adverse events in M patients, we shut down the trial for toxicity. IFF we have X or more complete responses in Y patients, we shut down the trial because it would be unethical to deprive the control arm. IFF we have Z or fewer responses in the treatment arm, given our ultimate accrual goal (total sample size), it will be impossible to conclude (using the test we have selected and preregistered) that the new drug isn't WORSE than the standard, so we'll shut it down for futility. Those patients will be better served by another trial. You are massively oversimplifying a well-understood problem. Decision theory is a thing, and it's been a thing for 100 years. Instead of lighting your strawman on fire, how about reframing it? Stopping isn't "always" wrong, but stopping because you've managed to hit some extremal value is pretty much always biased. The "Winner's curse", regression to the mean, all of these things happen because people forget about sampling variability. It's also why point estimates (even test statistics) rather than posterior distributions are misleading. If you're going to stop at an uncertain time or for unspecified reasons, you need to include the "slop" in your estimates. "We estimate that the new page is 2x (95% CI, 1.0001x-10x) more likely to result in a conversion"... hey, you stopped early and at least you're being honest about it... but if we leave out the uncertainty then it's just misleading. All of the above is taken into account when designing trials because not only do we not like killing people, we don't like going to jail for stupid avoidable mistakes.
- deleted 10y ago[deleted]
- tedsanders 10y agoMy point is that you can still extract useful information when your stop is dynamic rather than static. One typical scenario is when your effect size ends up being larger than you originally guessed. There's little reason to continue if the difference becomes obvious. In the future, I would appreciate it if you steelmanned my comments or asked for clarification instead of insulting me. It hurt my feelings. I wish I had written a better comment that hadn't incited such a reaction from you. Best wishes.
- imh 10y agoIf you decide to stop when a test says the error rate (FPR) is 5%, your error rate will be higher than 5%. If you don't want to call that a failure mode, it's at least a misuse of the metric.
- sogen 10y agoYes, stopping at 95% percent is BAD. First of all you need to reach on a sample size large enough, otherwise you are just lying to yourself. Source: I have a degree in analytics.