3 ms·
Arguably the "best" confidence interval for this situation is the Blyth-Still-Casella interval (preferred by StatXact), and the "best" hypothesis test is Barnar
by keithwinstein 13y ago
Arguably the "best" confidence interval for this situation is the Blyth-Still-Casella interval (preferred by StatXact), and the "best" hypothesis test is Barnard's test.
Here is code to calculate both: https://github.com/keithw/biostat https://github.com/keithw/biostat
I say "arguably" literally -- there is a huge body of literature on confidence intervals for binary proportions, much of it in disagreement about what is important. The Wilson score interval and Agresti-Caffo and whatever else are fine approximate methods that came of age when ease of calculation was a big concern. But if you have a computer and you're baking one thing into a library, may as well make it the best one you can.
Of course there is also plenty of merit to just picking some prior distribution and integrating over the conditional probability distribution given the data, aka a Bayesian approach.
In practice I don't think this (stats geekery about the merits of different confidence or credible intervals) is the most important part. The numerical results from all these techniques will be pretty similar.
The important part is in the design of the experiment, the interpretation, and playing by the rules. If you want to dynamically tune a Web site to make the most money as more information rolls in, that calls for a different experiment than a standard hypothesis test. (Even if you want to peek early at the results and possibly abort the test as a result, that calls for different tools and different rules.)
- noelwelsh 13y agoHadn't heard of the Blyth-Still-Casella interval or Barnard's test before, so thanks for that. By "If you want to dynamically tune a Web site to make the most money as more information rolls in, that calls for a different experiment than a standard hypothesis test." you're talking about minimising regret / bandit algorithms?