6 ms·
Winsorized mean
- laughy 3y agoA better alternative is to assume t-distributed errors
- rendywijaya 3y ago[flagged]
- SubiculumCode 3y agoIt can be useful data cleaning method when used judiciously, but I'm surprised its at the top of HN
- samvher 3y agoI thought I was pretty statistically literate, but learned a couple of new things here in the discussion, and this was not a tool that I would have thought of using before (now, in certain situations, I might think of it).
- nubinetwork 3y agoI would argue that random Wikipedia articles with no context shouldn't make it to the top of HN, but here we are...
- nlitened 3y agoPersonally, I love top-of-HN Wikipedia article submissions. Often an entrance to an interesting rabbit hole and thoughtful discussion.
- jedberg 3y agoThe trimmed mean and Winsorized mean are both super useful in monitoring and metrics systems. In most cases you don't actually want the median but you also don't want the extreme outliers to throw everything off with a mean. Both methods give you better metrics for comparing periodically, like day over day or week over week.
- PredictorX1 3y agoI've seen published estimates of statistical efficiency for a few test distributions for the trimmed mean, but not the Winsorized mean. Can anyone suggest references on this?
- deleted 3y ago[deleted]
- nivekkevin 3y agop99 for the same reason has been used widely in monitoring systems and benchmarks.
- bertil 3y agoFor instance, all commercial A/B testing services (as far as I can tell) compute the 99%-Winsorized stats for their reports. This is partially due to several common cases that can mess up an average computation, especially between two halves: How many times has the developer testing account clicked on “Buy”? Obviously, those should not been the dataset in the first place, but when dealing with corporate clients, the commercial thing to do is to fix the problem and not confuse the client with the details, unless they dig in the documentation.
- thih9 3y agoIs all fun and games until one or two of the extreme remaining values (that are later used to replace the rest) turn out to be an outlier in itself.
- munificent 3y agoYes, but by definition they must strictly be less of an outlier than the values they are replacing.
- gnulinux 3y agoRight but it can still pull the "Windsorized mean" to a particular direction. E.g. if your distribution is -99, 1, 1, 1, 1, 1, 2, 2, 2, 98, 99 and you "Windsorize" your mean, you'll pull it towards one direction. Here, we get: mean: 9.91 median: 1 Windsorized mean: 18.91 truncated (Olympic) mean: 12.11 2-truncated mean[1]: 1.43 as you see "Windsorized mean" is exaggerated on the positive direction. [1] this is when I remove top and bottom 2 values and calculate the mean of all the remaining values.
- jedberg 3y agoWindsorized (and trimmed) means really only work if you have one long tail, not two.
- pwdisswordfishc 3y agoA Windsorized mean is when you replace the outliers with members of the British royal family.
- gnulinux 3y agoOh whoops sorry about the typo!
- snicker7 3y agoWhy would I prefer a winsorized mean over a median?
- Scene_Cast2 3y agoAny of: low amount of data points; wanting a continuous value if your points are integers; slightly different behavior in high dimensional vector spaces.
- ivanbakel 3y agoFor a one-sided, unbounded distribution when you still want to observe changes without being susceptible to outliers. If you're monitoring response timings on a server, for example, the median might be very close to 0, and it won't shift unless a majority of the distribution slows down. If you take a winsorised mean, you can trim useless long response times that mess with the mean, but still see if e.g. 1/3 of your responses are suddenly slower than normal.
- timeagain 3y agoBut wouldn’t the trimmed mean mentioned in the article do this without “windsoriszing”?
- myhf 3y agoThe trimmed mean discards outliers, so it can measure "what is the typical value for non-outliers?". The windsorized mean reduces the weight of outliers, so it can measure "how many outliers are there?" instead of "how extreme are the outliers?"
- teamonkey 3y agoWhen processing astrophotos, multiple exposures are “stacked” together. There is a certain amount of noise in each frame - due to electronic noise or simply the random number of photons that strike a pixel in any one exposure - that you want to average out on a per-pixel basis. However some frames may contain unwanted outliers, for example if a satellite briefly passes overhead it will appear as a very bright streak in only one frame. By winsorizing, outlying pixel values can be eliminated while still maintaining the same number of samples per pixel as the rest of the stacked image.
- Alligaturtle 3y agoI'm not sure I would feel comfortable using the Winsorized mean -- it doesn't have any particular statistical properties, and it lacks any intuition appeal because it's not clear what the value represents. I can understand a line of logic that would give rise to something like the Winsorized mean -- after you look at your data, you see some obvious outliers. It feels dirty to just drop those values (which would lead to the truncated mean) because the information from an implausible value is more likely to be near the extreme than it is to be near the center mass. What to do with those extreme values? Here's something I now want to experiment with -- bootstrapping the extreme values. Take note of the original empirical distribution. Then, create a new distribution by removing the top and bottom X% of the observations and replacing them with values drawn i.i.d. from the original empirical distribution. This could lead to some values being replaced with the outliers that we originally wanted to drop. After we do this, record the mean. Then create new sample distributions until we have a distribution of new means. What I am curious about is how the shape if this distribution of means will be impacted depending on that X% value selected at the beginning. What are some well-known distributions that appear to have outliers? A log-normal distribution maybe?
- stdbrouw 3y ago> What are some well-known distributions that appear to have outliers? A log-normal distribution maybe? All of them. Which is why outside of a handful of contexts, the consensus in statistical modeling nowadays seems to be not to worry about outliers unless the values are completely unreasonable or there are a suspicious amount of them. As to your bootstrapping idea, why not bootstrap the entire distribution, why only the tails? If you only bootstrap the tails but allow draws from the entire empirical distribution then you are changing the underlying distribution.
- creer 3y agoDoesn't it depend on what is the likely reason for the outliers? - A world with a different distribution than the one you are trying to fit - A measurement environment subject to bad contacts or noise spikes or experimental mistakes - A reporting system with occasional typos - etc Seems to me what to do with the outliers should be informed by some understanding of the environment. And in some case, noted aside while waiting to see if there is more data "out there" in the outliers' vicinity. In some cases, a replacement for the outlier might be "nearby" while in other cases we know nothing about where the replacement should be.
- some_random 3y agoOne of the fascinating things about statistics is that in some ways it's more an art than a science, and the question of "when would I choose to use this over a normal mean or median" is a great example of that.
- phlip9 3y agoA related example out in the wild: Rust's `cargo bench` "winzorizes" the benchmark samples before computing summary statistics (incl. the mean). https://github.com/rust-lang/rust/blob/master/library/test/src/bench.rs#L150 https://github.com/rust-lang/rust/blob/master/library/test/s...
- croisillon 3y agonot to be confused with Florida mean
- bluenose69 3y agoI use it mainly as a data-exploration method. For example, if the Winsorized mean gives a value that differs a lot from the conventional mean, then I might examine the outliers in a bit more detail with tools like a boxplot or a histogram. The source of the data matters a lot in what methods make sense. For example, hand-entered numbers might involve transposed digits, or missing signs, or decimal points in the wrong place. Numbers deriving from some electronic measurements might have problems with numbers "pegging out" at some limit. In other cases, those numbers might "wrap around". Data that have been examined at an earlier stage might have numbers changed to something that is obviously wrong, like a temperature of -999.999 or something. The list goes on. My point is that exploring outliers is often quite productive, and comparing means to Winsorized means can be a very quick way to see if outliers are an issue. This is not so much an issue for interactive work, for which plotting data is usually an early step, but it can come in handy during a preliminary stage of processing large datasets non-interactively. It can also be handy as part of a quality-control pipeline in a data stream.
- fsckboy 3y agowhen you talk about Winsorizing, you're talking about applying some subtle sophistication to statistically sparse data... so then you mention histograms, which also apply subtle sophistication, so you best be talking about histograms with unequal bar widths. Equal bar widths is a bar chart. It doesn't get to be called a histogram (pace wikipedia) unless every bar reflects the same proportion of the population.
- stdbrouw 3y agoIf every bar reflects the same proportion of the population then that would mean every bar of the histogram would have exactly the same height, because the y axis on a histogram is density or proportion. I don't get how you can be so confident about a concept you only vaguely remember from a class you took long ago, and that even after reading the relevant article on wikipedia you conclude that therefore wikipedia must be wrong rather than thinking "hmm, wonder whether I'm misremembering things here..."
- Bjartr 3y ago
- hornban 3y agoThere's some interesting discussion in this thread about truncated vs winsorized means. For my own part, this is the first time I've come across either of these terms. I tend to benefit the most from seeing the entire distribution visually, and that helps me decide if I'm looking for a median, a "normal" mean, a "mean minus some weird outliers", or something different entirely. Does anybody happen to know of a good visual guide for how different measures of central tendency apply to various distributions? Anything that emphasizes pathological cases is helpful.