3 ms·
When I started learning about Bayesian statistics years ago, I was fascinated by the idea that a statistical procedure might take some data in a form like "94%
by bbminner 3y ago
When I started learning about Bayesian statistics years ago, I was fascinated by the idea that a statistical procedure might take some data in a form like "94% positive out of 85,193 reviews, 98% positive out of 20,785 reviews, 99% positive out of 840 reviews" and give you an objective estimate of who is a more reliable seller. Unfortunately, over time, it become clear that a magic bullet does not exist, and in order for it to give you some estimate of who is a better seller, YOU have to provide it with a rule for how to discount positive reviews based on their count (in a form of a prior). And if you try to cheat by encoding "I don't really have a good idea of know how important the number of reviews is", the statistical procedure will (unsurprisingly) respond with "in that case, I don't know really how to re-rank them" :(
- MontyCarloHall 3y agoWith that many reviews, any reasonable prior would have an infinitesimally small effect on the posterior. Assuming the Bernoulli model in the blog post, the posterior on the fraction f of good reviews is proportional to p(f|good = 85k*0.94, bad = 85k*0.06) ∝ f^79900*(1 - f)^5100 * f^<p_good = prior number of good review> * (1 - f)^<p_bad = prior number of bad reviews> ∝ f^(79900 + p_good)*(1 - f)^(5100 + p_bad) Recall that the prior parameters mean that the confidence of your prior knowledge of f is equivalent to having observed p_good positive reviews and p_bad negative reviews. So unless the prior parameters are unreasonably strong (>>1000), any choice of p_good and p_bad will have negligible effect on the posterior. The main reason Bayesian statistics is not a magic bullet is because it's up to you to interpret the posterior distribution. What really does it mean that the fraction of positive reviews from seller A is greater than the fraction for seller B with probability 0.713? What if it were 0.64? 0.93? That's for you to decide.
- nextos 3y agoYes, two extra notes on this: 1. Weakly informative priors can be good to regularize and stabilize inference in scenarios with a low amount of data. 2. In case of the review scenario presented by the OP, a hierarchical model could be even better as it would achieve regularization (and shrinking) by borrowing information across different sellers. In other words, one would learn the overall distribution of seller reliability at the same time as individual ones. This has two advantages: a) Just one (hyper)prior for the overall distribution, but no priors needed at seller level and b) seller predictions are pulled towards the mean in a principled way. Most scientific publications that find unusually large effects coming from some association turn out to be false. If they used shrinking (2), they would avoid excessively optimistic predictions. See: https://en.wikipedia.org/wiki/Stein%27s_example https://en.wikipedia.org/wiki/Stein%27s_example
- TeMPOraL 3y agoAnd then there's real life: it's better to look at number of annulled negative reviews, if that stat is published. Positive reviews are gamed as standard practice. Negative reviews are bribed away. But the number of negatives removed is a good proxy for how bad the seller is.