5 ms·
Does anyone have an explanation of PCA that's more accessible to laypeople / the less mathematically inclined?
by hammeiam 8y ago
Does anyone have an explanation of PCA that's more accessible to laypeople / the less mathematically inclined?
- salty_biscuits 8y agoWhich way does the blob point
- salty_biscuits 8y agoSlightly less tersely, if you fit a multivariate gaussian to some points you could think of this as an something like an ellipse (at a contour of constant probability). The semi major/minor axes of the ellipse are the principal component directions. The size of the principal components is like the radii of the ellipse.
- whytaka 8y agoI just read this today. You may find it helpful. http://setosa.io/ev/principal-component-analysis/ http://setosa.io/ev/principal-component-analysis/
- pepper_sauce 8y agoA similar example (using the popular Iris dataset): https://www.math.umd.edu/~petersd/666/html/iris_pca.html https://www.math.umd.edu/~petersd/666/html/iris_pca.html
- Eli_P 8y agoIt takes a cloud of points as input, and gives a set of shapes that approximate that cloud, noise filtered out. So component is one of those shapes. In case of PCA, shapes are ellipses. For SOM, shapes could be more complex like zigzag. For k-means, shapes would be like Voronoi cells.
- carlmr 8y agoYou have certain values that you use to predict another value. Now you have maybe a lot of criteria, some more and some less important ones. You try to find correlations here between the inputs, so that you can describe the output more simply with a combination of these input values.
- mikorym 8y agoBy way of example, suppose you have a weather dataset with 30 parameters such as temperature, cloud cover, solar radiation, etc. Your objective is to classify points on a map that have similar weather conditions. You can do PCA on these geographical points at a given time. Your matrix will have as columns weather parameters and as rows the different geographical points. PCA will determine how to best represent the variation in data into smaller dimensions. If you are lucky, 80% of the variation can be plotted in a 2D plot as in OP's post. Points in the x/y cartesian that have similar weather will plot closer to each other. This process reduces dimension. For example, temperature and solar radiation are strongly correlated and will tend to "push" into the same direction so that you can replace those parameters by a virtual parameter. (For weather data, that would tend to be the x-axis in the PCA biplot as temperature and similar parameters are often sufficient to classify a climate for the purposes that I have encountered in agriculture.) In a nutshell, this is a real word example of when one can use PCA.
- siddboots 8y agoA key point that I feel is missing from this explanation is that the concept of distance becomes almost meaningless in very high dimensional space [1]. PCA is a means of dimensionality reduction, but the reason you want to reduce dimensions is so that you can measure and compare the distance between pairs of samples. [1] https://en.m.wikipedia.org/wiki/Clustering_high-dimensional_data https://en.m.wikipedia.org/wiki/Clustering_high-dimensional_...
- yorwba 8y agoThe "meaninglessness" of distances in high-dimensional spaces, as defined in that Wikipedia article (i.e. the maximum and minimum distances of points becoming relatively close to each other) only happens when all those dimensions are independent. If the data is actually low-dimensional, such that PCA loses no information (trivial case: tack a billion zeros to all your vectors) then it won't change the distances at all, so it doesn't matter whether you compute them in the high-dimensional space or the low-dimensional projection obtained with PCA. PCA works under the assumption that you're dealing with a low-dimensional signal that was projected into a high-dimensional space and then corrupted by low-magnitude high-dimensional noise. By taking only the top k components, you filter out the noise and get a better signal. But if you have a low-magnitude high-dimensional signal corrupted by highly-correlated noise, then naively applying PCA will filter out the signal and leave you with only noise. So whether PCA makes your data more meaningful or less really depends on whether that assumption is satisfied or not. Dimensionality reduction is no silver bullet.
- masthead 8y agohttps://stats.stackexchange.com/questions/2691/making-sense-of-principal-component-analysis-eigenvectors-eigenvalues https://stats.stackexchange.com/questions/2691/making-sense-...
- IngoBlechschmid 8y agoHere is a neat demo which uses PCA for image compression: https://timbaumann.info/svd-image-compression-demo/ https://timbaumann.info/svd-image-compression-demo/
- j7ake 8y agoPCA fits a high-dimensional ellipse to a cloud of points. The axes of the ellipse correspond to the eigenvectors. The length of each axes correspond to the eigenvalues. You can reduce dimensions by saying some of the eigenvalues are "noise" and discarding them.
- alexcnwy 8y agoI found this explanation extremely helpful when I did PCA in my postgrad multivariate analysis course: http://setosa.io/ev/principal-component-analysis/ http://setosa.io/ev/principal-component-analysis/
- kk58 8y agoLet's take the terms,first :prinicipal , second :component , third :analysis. Principal is kinda mathematic's way of saying something is important or critical. Component is basically referring to parts Analysis is basically referring to ability of this method to help you understand what's going on with your data. So the main idea here is like actually pretty simple. Let's say you want to understand what makes a difference to a person's SAT score. You track number of hours of study, maybe the average school SAT scores, age of the person, average of mock SAT scores. Now you know that some of the columns of data are more important than other and you want to know which one? So in regression what you would do is fit a line that best lies in the middle of these points by playing around weights or 'importance' of the columns till you get least distance from points. What PCA does is it tells you which variables are important and gives you a sense of ranking of these variables. So it has the ability to select variables, rank it's importance and tell you what columns of data matter more than others. Ergo it tells you which components are principal and helps you analyse them through their importance ranking. Because this method does so many things you can a do a ton of cool stuff. If you have a ton of data, this method can tell you to focus on these 3 or 4 columns which have the biggest impact. So it can help you prioritize Second if you are looking at optimizing your system,let's say SAT scores, this system can tell you a better school can make a bigger difference than just brutal hours of practice. In networks like social networks, it can tell you who is the most important/ prestigious/ coolest person by looking at friendship or social messaging links between people. So to sum up, it gives you an idea of what is important in your data, gives a sense of the quantum of its importance and hence gives a deeper feel for what's going on. One big point with PCA is that it's a linear method. Which means variables which have exponential impact on your study will not get signalled well. So transformation and processing your data is critical for this method to work. Hope this helped.
- oktavist 8y agoDraw a banana on a piece of paper. Chances are in doing so you subconsciously eliminated the dimension of the banana that is 'non curvy' / 'not really varying as much'. You've mentally done a form of PCA, and eliminated the least interesting dimension in reducing it to 2 dimensions instead of 3. It is a means of defining a new coordinate system that is ordered by dimensions of decreasing variance.