6 ms·
> The current state of the art of ML assumes nonlinear relationships between all parameters. It can't assume simpler & reasonable models, and therefore it can't
by throwawayjava 8y ago
> The current state of the art of ML assumes nonlinear relationships between all parameters. It can't assume simpler & reasonable models, and therefore it can't extrapolate easily with reduced data.
I'm not really sure what "low-dimensional intuition" means, but I pretty regularly build models that do not "assume nonlinear relationships between all parameters".
- jacques_chester 8y agoMy understanding is: It's a fancy way of saying "lots of variables", because a lot of these problems are converted into linear algebra representations (which can have geometric analogies, hence: dimensions).
- toufiqbarhamov 8y agoI think the Wikipedia article on the Curse covers it pretty well. https://en.m.wikipedia.org/wiki/Curse_of_dimensionality https://en.m.wikipedia.org/wiki/Curse_of_dimensionality The common theme of these problems is that when the dimensionality increases, the volume of the space increases so fast that the available data become sparse. This sparsity is problematic for any method that requires statistical significance. In order to obtain a statistically sound and reliable result, the amount of data needed to support the result often grows exponentially with the dimensionality. Also, organizing and searching data often relies on detecting areas where objects form groups with similar properties; in high dimensional data, however, all objects appear to be sparse and dissimilar in many ways, which prevents common data organization strategies from being efficient.
- tomnipotent 8y agoMeaning that you can run PCA and the first few components will explain a vast majority of the variance. Some things may have so many moving pieces that it's just not possible to grok how all the variables interact to get the results we see.
- gnulinux 8y agoAs the dimension d goes to infinity, if you sample integers from d-multivariate normal distribution, your samples will collect near the unit d-sphere; whereas in low dimensions they'd collect near the origin. This is because as d goes to infinity, your samples become more and more far apart. This is a topological intuition of why higher dimensions don't work like lower dimensions.
- throwawaymath 8y agoBeautifully and concisely explained!
- DoctorOetker 8y agoI agree it's beautiful and concise, but it is not correctly explained. I can give you a beautiful and concise explanation of clouds too: as we all know objects fall back to earth, or more properly speaking they follow orbits, close to the surface they seem to follow parabolas, but in fact they follow ellipses. the moon orbits the earth in a nearly circular ellipse, as do geosynchroneous satellites. When water evaporates from the ocean they are in fact launched into an ellipcitcal orbit, however each elliptical orbit that intersects the surface of the earth will intersect it again, this explains rain, the individual water drops fly for a while until they fall back down... beautiful, and relatively concise in comparison with real cloud physics, but totally wrong of course...
- DoctorOetker 8y agoI reserve the possibility that I am completely mistaken, in which case I apologize in advance, on the condition that you refer me to an actual calculation or derivation (not just a quote in the same spirit) in a paper or text or textbook. I have seen such insinuation multiple times, that samples "collect" near the unit d-sphere when samples are drawn from the unit d-sphere in high d dimensions. From a physics perspective this is very familiar to me, but not in the sense of fact, but in the sense of misinterpretation. I do believe the observation is very useful in the educational sense as long as it is pointed out as being a paradoxical illusion. In this sense I can appreciate (even encourage) a professor or a TA showing this phenomenon, on the condition they finish up by explaining why this seems to be the case, but is nevertheless a misinterpretation, and they should make sure the connection with Jacobian determinants etc are made clear. Consider a normal distribution of any dimension (as high as you want), but I will showcase the phenomenon even with low dimensions (here merely d=3) to illustrate this has nothing to do with high dimensions. Clearly the probability density is maximal in the center of the distribution. In computer processing of data points, we typically loop over points, calculate some hopefully interesting function on each sample, and then plot the samples say by binning with equal bin sizes. A programmer typically disregards transformation properties like the Jacobian determinant. Suppose the value we calculate for each data point is the absolute length or distance from the center. The further we go from the center the smaller the probability density of the normal distribution becomes... but the larger the volume of a shell of radius r! Since we are binning with equal bin sizes (equal length intervals per bin), for small lengths, then even though the actual probability density is highest near the center, we will get relatively few samples because the volume under consideration is small compared to the volume under consideration for a shell of a larger radius of equal thickness (area of a sphere grows quadratically with radius). However for even larger distances, the exponential decay of the normal distribution dominates and the number of samples in highest radius bins will decrease again. So in between there will be a peak. This explains the fact that [ the probability density of [ the absolute distance from the center over [ the sample points ] ] has a peak at some non-zero length. But it is a conceptual mistake to interpret this as if those sample points in the original d-dimensional space form a dense shell on some "unit sphere" ... This is a pure illusion which illustrates the interpreter is not familiar with jacobian determinants etcetera. Consider volume or triple integrals over some volume element dx dy dz and for symmetry you prefer integrating in a spherically symmetric coordinate system, theta,phi,r then you can not simply replace dx dy dz with dtheta dphi dr, you need to use dV = dx dy dz = r^2 sin( phi ) dtheta dphi dr. It is this necessary factor that is ignored when processing sample by sample and causing this illusion in the AI community. I did not follow conventional machine learning courses, but given the learned language used whenever I see statements to the effect of samples lying near the unit sphere in high dimensions, I can only conclude it has its origins in 1) direct observation or experience of plotting in bins of the length of the vector without guidance in interpretation; or 2) guidance during education, where the phenomenon is shown, and the origin of the paradoxical illusion or confusion adequately explained and then the illusory nature subsequently forgotten or 3) a teachers assistant having gone through 2) and showing the phenomenon to students without emphasizing the illusory nature of the misinterpretation. But perhaps I am wrong, and the normal distribution in high dimensions actually has a higher probability density near its "unit sphere" if the dimension is beyond some critical dimension d_c... but again, I'd like to see a derivation showing it :) EDIT: For more precise language: they do "collect" (reach a peak) at a certain non-zero length or distance, but they do not collect to a unit sphere of such radius in the original space!