4 ms·
Couple comments: Anyone in statistics/probability has seen this quote multiple times: "All models are wrong, but some are useful." -George Box. We often don't
by christopheraden 13y ago
Couple comments:
Anyone in statistics/probability has seen this quote multiple times: "All models are wrong, but some are useful." -George Box.
We often don't know the underlying distribution of a particular process or our data, true, but that doesn't mean that we can't get useful information out of it by making some assumptions, provided that our assumptions are not wildly wrong. For example, if we assume height is normally distributed, that should mean it's possible, under our model, to see someone who is -9000 feet tall. We're willing to accept this falsity in exchange for having some nice properties we wouldn't have had otherwise, provided we exercise some common sense.
In regards to (3), if all you can wield is the SLLN, you can't do much inference or prediction. So you're able to guarantee that the sample mean will converge to the expectation almost surely. The closeness of approximation in a finite sample case using asymptotics is intrinsically tied to sample size and the particulars of the underlying distribution of the random variable.
In regards to (5), the curse of dimensionality does become a large problem, but throwing out assumptions makes it a far bigger problem. The nonparametrics community has a lot of trouble with multivariate distributions for precisely this reason! This is one area where parametric models do better than non-parametric ones, since they have so much more structure to play with, it makes the problems so much more tractable.
You mention order statistics... order statistics in higher dimensions is a tricky concept that is not as defined as the one dimensional case. Which one is more: (3, 2, 1) or (1, 2, 3)? Wouldn't that necessarily depend upon the underlying distribution or some other measure of distance? If you have a paper to suggest that general high-dimensional order statistics makes sense, throw it this way. You are correct that Least Squares doesn't require a distribution to get mean and variance estimates, but what good is that if you don't know how much of your population fits within a multiple of your standard deviation around your mean? You could use something that works for all distributions like Chebyshev's Inequality, but to get that level of sweeping generality seriously hurts your power (If your data IS actually normal, 95% of obs fall within 1.96 SD's of the mean. Chebyshev will tell you 75%), and even Chebyshev had to impose a restriction to get that bound--it assumes finite mean and variance.
While there are distribution-free hypothesis testing methods, most of classical nonparametric statistics makes some assumptions about the underlying distributions, albeit they are far gentler than the parametric assumptions.
Always choosing nonparametric methods is just as short-sighted as always using parametric methods. It's very possible that you are throwing away vast amounts of information by using a nonparametric procedure when a parametric one would have been completely adequate. With ANOVA and two-sample tests, you don't lose much by going with Kruskall-Wallis or Mann-Whitney tests (under normality, MW has a relative efficiency of 3/pi versus the t-test), but in other circumstances, using a nonparametric method could be way worse, provided the parametric assumptions are true.
What I'm getting at is that while non-parametrics is nice in that it doesn't require many assumptions, they may be throwing away too much. Fitting to a distribution may very useful in making inferences and predictions. All models are wrong, but some are useful.
- graycat 13y agoSpoken like a true statistician! Yes, the statisticians keep assuming Gaussian, fitting distributions to data, etc., and you have a way out: Of course it's wrong, but it's still useful! Wow! > If you have a paper to suggest that general high-dimensional order statistics makes sense, throw it this way. Do an hypothesis test. For positive integers m and n, consider m > 1 points in real Euclidean n-space. We want to know if point m is distributed like the rest. Our null hypothesis is that all the points have the same distribution and are i.i.d. Now for positive integer k, for each of the m points, we can find the distance to the farthest of the k nearest neighbors. So, we get m distances. These, however, are not i.i.d. But can do a little work to show have a finite group of measure preserving transformations such that, if we sum over the group, then we get something similar to a permutation test. So, if the distance for point m is in the upper 2% of the distances, then have probability of type I error of 2% and, thus, an hypothesis test. So, have a multidimensional, distribution-free hypothesis test. That is, intuitively if point m is 'too far away' from the other m - 1 points, then we reject the null hypothesis that point m is distributed like the other m - 1 points (we continue to believe in the independence assumption and do not reject it). But the k-nearest neighbors with the Euclidean norm need not be the only one used -- so get a family of such tests. Intuitively, for n = 2, we have a density that looks like, say, a mountain range. Suppose the rocks of the range are porous to water and pour in some water. Then the water all seeks the same level. So, we have multiple lakes, with islands, with lakes, etc. Suppose the water covers 2% of the probability mass. Then the test is to see if point m falls in the water. Yes, we expect the lake boundaries to be fractals. So, right, with enough points m, we are approximating the fractal boundaries of the lakes. Increasing k makes the approximation to the boundaries more smooth. If pour in more water, then we get another set of contours. So, we get a way to do contour maps of the density. So, we get a technique of multidimensional density estimation. E.g., go to a beach, pick up m = 200 rocks, and for each measure n = 10 properties. Now given one more rock, ask if it came from that beach! As I recall, there is such a paper in 'Information Sciences' in 1999 about an "anomaly detector'. Why 'Information Sciences'? Because the paper suggested monitoring large server farms and networks for 'health and wellness' by this technique. So, we would get multidimensional, distribution-free 'behavioral monitoring' with false alarm rate we could set in advance and get exactly. The work does not promise to have the best power in the sense of the Neyman-Pearson result, but the paper uses Ulam's 'tightness' to argue that the test is not 'trivial'. I have yet to see a distribution-free test that tries to argue it is as powerful as Neyman-Pearson! While we don't get Neyman-Pearson, we get an approximation to the smallest region where we will make a type II error and, thus, in a sense for any alternative distribution for point m the most power in the goofy sense of shifting that alternative distribution around! The argument is just a simple use of Fubini's theorem! So, in a goofy but possibly "useful" sense we minimize type II error. There might be a duality situation here!