5 ms·
Why? They’re non-parametric and make zero assumptions of normality.
by pocketsand 2y ago
Why? They’re non-parametric and make zero assumptions of normality.
- blueflow 2y agoHow else would you calculate the quartiles to render the boxes?
- munch117 2y agoCount data points in each quartile. You can do that for any sortable data, independent of distribution.
- blueflow 2y agoIf you do that in your paper, you better write next to the graph that you did that.
- thaumasiotes 2y agoArguing that nobody who might be professionally expected to look at a box plot can be reasonably expected to understand how box plots are defined doesn't make a compelling case that using them is a good idea.
- blueflow 2y agoIf the method how the plot boxes are calculated is not clear (this thread references at least two different methods), you'll need to explicitly write it down which methods you did use.
- thaumasiotes 2y ago> this thread references at least two different methods No, as the sidethread comment notes, there is only one way you can compute quartiles. You seem to be arguing that the correct thing to do is to impute them, and that calculating them is such a deviant practice that it would need to be specially remarked on.
- blueflow 2y agoIsn't this what i was saying from the beginning? Box plots are made for visualizing generalized normal distributions and nothing else. And now people in this thread argue you can calculate them from something else. Not sure if you are replying to the right post.
- thaumasiotes 2y agoThat might be what you were saying from the beginning, but the only thing that that would establish is that you're completely out of touch with reality. Box plots are made for visualizing quartiles. Your theory would imply, among other things, that the median line going through the box part of a box plot always divides it in half, which obviously is not the case.
- blueflow 2y agoNo? Exponential Gaussian? Whatever you do, you should explain first what you do that your whiskers stay meaningful and are not just whatever randomness your outliers produced.
- A4ET8a8uTh0 2y agoIt is actually a fascinating argument that shows how little of what is being decided is based on actual data ( or at least our understanding of it ), but rather that data visualization is being used to push already pre-approved decisions with data being used merely as a 'for' argument. I agree that if there is an indication that if most professionals don't really know what boxplot is supposed communicate, maybe it should not be used.
- munch117 2y agoPerhaps I expressed myself poorly, and left room for misunderstanding, because I cannot possible imagine that we have any real disagreement on how to compute quartiles. Any set of numbers I give you, you can compute quartiles for it. There is no algorithm for doing that that breaks down if the numbers don't follow a normal distribution.
- blueflow 2y agoLook at this SVG from wikipedia: https://upload.wikimedia.org/wikipedia/commons/1/1a/Boxplot_vs_PDF.svg https://upload.wikimedia.org/wikipedia/commons/1/1a/Boxplot_... When you calculate the box plot using normal distribution parameters, the outliers are outside the outer bracket. If you split the dataset into 4 equal parts, the bracket will be larger because the outliers are still inside it. The methodologies are not equal. This thread is the first time i heard people do the "split dataset into 4 quarters" and using that for box plots.
- pocketsand 2y agoAs I'm sure you know, there are a lot of variations on how quantiles are calculated in various software. The 25th percentile, e.g., doesn't always line up with a value in the dataset, so sometimes nearest rank methods are used, otherwise a linearly interpolated data point, where interpolation is done in various ways. In any event, none of these methods assume normality, or rely on CDFs of a normal curve. If they did, every box plot would be symmetric. The fact some people think that boxplots are constructed in such a way is a pretty good reason to take the author's article seriously as for how boxplots are confusing.
- ColFrancis 2y agoAs a first pass definition it does well to explain the concept. Even if you're interpolating you will need to rank the samples and find the two nearest neighbours to interpolate between. It serves to distance it from the moment-based statistics like mean and variance at least.
- 2y ago
- blueflow 2y agoOn second thought, this method makes the outer brackets / whiskers pretty much useless since their position is determined by the largest outliers, which is quite much random.
- Falkon1313 2y agoThat's not how they're drawn. Outliers (More than 1.5 times the interquartile range outside the 1st/3rd quartile) are plotted as dots beyond the whiskers. The whiskers go at Q1-1.5×IQR and Q3+1.5×IQR.
- blueflow 2y agoBetter is! Look what i was replying to.