4 ms·
This is a good discussion, but the author is confused about what the "Low Base Rate Problem" is. It doesn't have anything to do with the null hypothesis being t
by EvanMiller 12y ago
This is a good discussion, but the author is confused about what the "Low Base Rate Problem" is. It doesn't have anything to do with the null hypothesis being true most of the time -- the example the author gives is actually a second form of repeated significance testing, which could be addressed with a Bonferroni or Šidák correction.
The Low Base Rate Problem is when you have a binary outcome and one of the outcomes is rare (say, less than 1%). There is so little entropy in the information source that you have to acquire a heck of a lot of samples in order for the statistical test to have any power. The problem is not unique to frequentist statistics; it's a consequence of information theory and so it affects Bayesian statistics as well.
Nonetheless, I highly recommend examining Bayesian test techniques to avoid repeated significance testing (both within a single trial and across multiple trials). A side benefit is that when someone says "What's the probability that the new purple dragon logo outperforms the old one?", you can give them an answer without backpedaling and explaining null hypotheses, p-values, significance levels, and all that jazz.
The major drawback to Bayesian techniques is that it tends to be computationally expensive. For example, to evaluate the A/B test and answer the purple-dragon question with normal priors, you have to integrate a normal distribution in two directions, and there's not a clean analytic formula for that. That's why there's a jagged histogram in the blog post; it changes every time you hit "Calculate" because it's being integrated with Monte Carlo techniques, which take a lot of juice compared to (frequentist) analytic methods.
- kobyszcze 12y agoThanks for the comment! What we were trying to get at is running repeated experiments when prior probability of an experiment being successful is low---which you correctly point out is also about repeated testing (and has nothing to do with power). So perhaps our naming is unfortunate.
- kiyoto 12y ago>A side benefit is that when someone says "What's the probability that the new purple dragon logo outperforms the old one?", you can give them an answer without backpedaling and explaining null hypotheses, p-values, significance levels, and all that jazz. This is surprisingly valuable. As a former quant/math person, I am shocked time and again how most self-professed data driven people have no idea about inferential statistics. I've learned over time that most people just want to see data and descriptive statistics, preferably in visual representations, and interpret them somewhat creatively and pretty much non-rigorously.