4 ms·
I'm not sure I follow this, or buy into the suggestion of the post. First, I don't see how it's true that data with a relatively large amount of variance will
by brockf 13y ago
I'm not sure I follow this, or buy into the suggestion of the post.
First, I don't see how it's true that data with a relatively large amount of variance will tend to be power law distributed. Defining what a "large amount" of variance is is tough (it depends on your intuition and choice of variance metric) but there are lots of distributions with considerable variance that are, for example, normally distributed (many more than are power law distributed, as far as I can tell).
Second, if you find that this is misleading your projections, why not just use a different kind of average? For example, if you just want to know, "How much is the next customer likely to spend?", you might use the mode. Or, if you want a more robust average (i.e., less likely to be seriously thrown off by outliers), why not use the median? You can even complement these with confidence intervals if you want to get a sense of their precision.
Like twic already said, you need some indicator to understand what's going on with your business. I think that in many cases, this will be the mean. But if you want something more robust or more practical, perhaps the median or mode might suit you better.
- brockf 13y agoI should also add... you can also use median absolute deviations, standard deviations, or interquartile ranges to identify and remove outliers who you think don't reflect your business's true status. But it all depends on what you want your model to do!
- raviparikh 13y agoTo address your first point – that's fair, I didn't really dive into the statistics in detail (since it wasn't really the point of the post). If variance of a distribution is finite then the central limit theorem applies, and given a sufficient number of trials, the distribution will begin to approach a normal distribution. However some data sets have infinite variance and may (under some conditions) begin to approach a power law distribution; or, they have finite but extremely large variance, in which case the number of trials it will take to begin to look like a normal distribution is very, very large. For your second point – there are legitimate uses for other summary statistics (mode, median), but they can still be very misleading. For example, you mention using the mode of the distribution as a predictor of what the next customer will pay – this definitely wouldn't work for a company with a metered pricing plan, for example. Distributions are often not well characterized by singular summary statistics, Anscombe's Quartet (mentioned elsewhere in this comments section) is a good example of this: http://en.wikipedia.org/wiki/Anscombe's_quartet http://en.wikipedia.org/wiki/Anscombe's_quartet
- brockf 13y agoRe: power law distributions. I'm not sure what you mean about finite versus infinite variance. Are you referring to whether you are analyzing a bounded versus unbounded scale? Even if a scale is unbounded, that still wouldn't make a dataset any more likely to power law distributed. A power law distribution would be, however, somewhat more likely to be observed in cases where there is only an upper- or lower-bound (though again, not always... it really depends). Re: appropriate averages. Your case (metered billing) is an interesting one. I don't see why the mode is necessarily wrong - the next customer is most likely to spend the amount that is currently your most popular amount). However, in order to calculate a mode, you likely would want to bin your amounts into ranges so that $10.11 isn't treated as distinct from $10.12, etc. You're definitely right about one summary statistic not being sufficient. I would advocate for visualization and summary statistics, with some estimate of your confidence in the estimate displayed visually.