3 ms·
It does, if you assume you care about the validity of the results or about making changes that improve your outcomes. The degree of care can be different in le
by epgui 1y ago
It does, if you assume you care about the validity of the results or about making changes that improve your outcomes.
The degree of care can be different in less critical contexts, but then you shouldn’t lie to yourself about how much you care.
- renjimen 1y agoBut there’s an opportunity cost that needs to be factored in when waiting for a stronger signal.
- Nevermark 1y agoOne solution is to gradually move instances to you most likely solution. But continue a percentage of A/B/n testing as well. This allows for a balancing of speed vs. certainty
- imachine1980_ 1y agodo you use any tool for this, or simply crunk up slightly the dial each day
- travisjungroth 1y agoThere are multi armed bandit algorithms for this. I don’t know the names of the public tools. This is especially useful for something where the value of the choice is front loaded, like headlines.
- hruk 1y agoWe've used this Python package to do this: https://github.com/bayesianbandits/bayesianbandits https://github.com/bayesianbandits/bayesianbandits
- epgui 1y agoEven if you have to be honest with yourself about how much you care about being right, there’s still a place for balancing priorities. Two things can be true at once. Sometimes someone just has to make imperfect decisions based on incomplete information, or make arbitrary judgment calls. And that’s totally fine… But it shouldn’t be confused with data-driven decisions. The two kinds of decisions need to happen. They can both happen honestly.
- scott_w 1y agoThere is but you can decide that up front. There’s tools that will show you how long it’ll take to get statistical significance. You can then decide if you want to wait that long or have a softer p-value.
- Jemaclus 1y agoI don't think I'm making the case that you shouldn't test things or care about the results, but rather a matter of degree of risk that should be acceptable. In medicine, if you get it wrong, people /die/. In software, if you get it wrong, /you sell fewer widgets/. That's a pretty major difference. You can't get it wrong in medicine, but you /can/ get it wrong in software without it being catastrophic failure. I'm basically making the case that "Your startup deserves the same rigor [as medical testing]" is making a pretty bold assertion, and that the reality is that most of us can get away with much less rigor and still get ahead in terms of improving our outcomes. In other words, it's still A/B testing if your p-value is 0.10 instead of 0.05. There's nothing magical about the 0.05 number. Most startups could probably get away with a 20% chance of being wrong on any particular test and still come out ahead. (Note: this assumes that the thing your testing is good science -- one thing we aren't talking about is how many tests are actually changing many variables at once and maybe that's not great!)