8 ms·
For readers who are OK with some math, I recommend John Myles White's eye-opening post about means, medians, and modes: http://www.johnmyleswhite.com/notebook/2
by andy_wrote 10y ago
For readers who are OK with some math, I recommend John Myles White's eye-opening post about means, medians, and modes: http://www.johnmyleswhite.com/notebook/2013/03/22/modes-medians-and-means-an-unifying-perspective/ http://www.johnmyleswhite.com/notebook/2013/03/22/modes-medi... He describes these summary descriptive stats in terms of what penalty function they minimize: mean minimizes L2, median minimizes L1, mode minimizes L0.
A single-number statistic is _going_ to leave things out, so if you must boil things down into one number, or even a few numbers, you're going to lose something that you had in the raw data. This is why I find claims along the lines of "statistics don't tell the whole story" a little bemusing - of course they don't, the very definition of a statistic is a summarization of data that is easier to work with. The question is what data is kept or lost, or more generally what importance we place on different aspects of the raw data such that it's reflected in our descriptive statistics.
The lessons for non-technical people who want to communicate with descriptive statistics are to recognize that summarization is inherent in the nature of any descriptive statistic, that they are thereby opinionated in some way in terms of what they've preserved and what they've left out, and to recognize whether those opinions are appropriate for your purpose.
- johnmyleswhite 10y agoGlad you enjoyed that post so much. It really is a shame that we do such a bad job of teaching students about the inherent subjectivity of descriptive statistics and let students leave their courses with dangerous ideas about the existence of a Holy Grail statistic that will solve all of their problems.
- nkkar 10y agoYour followup post (http://www.johnmyleswhite.com/notebook/2013/03/22/using-norms-to-understand-linear-regression/ http://www.johnmyleswhite.com/notebook/2013/03/22/using-norm...) is excellent. Thank you!
- johnmyleswhite 10y agoThanks! I really should have finished and written the post about the SVD as well. One of these days...
- agandy 10y agoI'd love to read your post on SVDs once it's written
- andy_wrote 10y agoThanks for writing it! It's one of my favorite math blog posts floating out there.
- dhfhduk 10y agoI agree with you, although I think the problem is a focus on procedures rather than principles in general. It took me a long time to realize that a principled reason for gaussian parametric distributions is the maximum entropy principle. Prior to that point, it had been presented as essentially arbitrary, even by established professors. There seems to be an assumption that theoretical statistics is "too hard", and as a result there's a middle ground that gets left out. I haven't taught general stats courses in awhile (although I've taught advanced ones), but if I did, might start with Bregman divergences, and work down in the manner of your blog post. I think there's an in-between that gets lost. You can teach principles without deriving long proofs of everything along the way. Students don't get taught the underling principles and philosophies to choose from, and I think this leads to the "holy grail" issues you're referring to. It seems to be changing a bit with new interest in Bayesian methods, but that's just the tip of the iceberg.
- ttub 10y agoThe idea of misinterpreting metrics is a very general idea and is not specific to statistics. Humans want to distill vast amounts of information to a more manageable amount, like for example a single number. Equity analysts look at accounting metrics, psychologists look at psychometrics test cores, doctors look at some function of blood pressure, etc etc. Any person with deductive and sceptical mental faculties in place, will recognize that these are all simplifications, and cannot be used to deliver a unified truth. Also, being aware of this has very little to do with being technical or not (for example, plenty of programmers only look at a CPU's clock speed to gauge performance). Anyhow, nice post.
- sp332 10y agoI always pull out Anscombe's Quartet https://en.wikipedia.org/wiki/Anscombe's_quartet https://en.wikipedia.org/wiki/Anscombe's_quartet The four datasets have the same mean, variance, and linear regression line, but are very different from one another.
- FabHK 10y agoGreat example, and mentioned in the article.
- abhgh 10y agoThis is a great and oft-forgotten point. I like to think all summary numbers are lossy, you are only free to pick your poison.
- CalChris 10y agoStatistics are reductionist! Well yes, that is their purpose.
- IndianAstronaut 10y agoMy statistics professor once told us that statistics are a shadow of the truth, not the actual truth itself.
- matt4077 10y agoWhat a caveman!
- lalaithion 10y agoOkay, so that leads to a very obvious (imo) question I haven't seen anyone ask: What happens when you minimize Ln, with n > 2? Why don't we use any of those?
- ssalazar 10y agoIncreasing n gives increasing weight to outliers, to the point where L-infinity is just the single maximum value. I would guess this is less useful when attempting to understand how existing data can make conclusions about future aggregate/typical cases vs. analyzing specifically the outliers.
- Sniffnoy 10y agoClarification: The l-infinity norm of a vector is the maximum absolute value of its coordinates. The analogue of the median and the mean that that gets you is the midpoint of the range.
- deleted 10y ago[deleted]