4 ms·
> 1) Does linearly combining my features make any sense? > 2) Can I think of an explanation for why the linearly combined features could have as simple a relat
by platz 5y ago
> 1) Does linearly combining my features make any sense?
> 2) Can I think of an explanation for why the linearly combined features could have as simple a relationship to the target as the original features?
the article provided a negative example where pca does not fit, but doesnt provide an example where it does or what pca is actually used for. I come away from this article thinking pca is useless.
what would be an example where 2) is true?
I cannot answer 2) without already having experience of what explanations there could possibly be. ( 2) is almost begging the question, at least pedagogically- pca is good when the features are good for pca)
When does linearly combining my features "make sense"? again, an example is not provided
- ThereIsNoWorry 5y agoI mean, that's the point. What kind of real world complexities are actually linearly explainable? Almost none. It goes all the way back to why classical statistics failed to provide real world value after a certain point and the trend has been going towards using non-linear black boxes for the last decade.
- platz 5y agoSo you basically are saying PCA is a useless procedure. Also, "after a certain point" is doing a lot of heavy lifting in that sentance.
- hervature 5y agoThe article is silly because PCA cannot select features. It is all about dimensionality reduction. You should think of PCA as the equivalent to VAEs from the neural network world. The idea would be something like this: you have big images (let's say 4k) and this is too expensive to train with/store forever. So, you collect a training set, train a PCA on these images, and then you can convert your 4k images to 720p or even 10 numbers, which you then use to predict/train whatever you want. Of course, we have algorithms that scale images but maybe all your images are of cats and there is a specific linear transformation that contains more information from the 4k image than simply scaling. The implicit thing here is that you still are collecting 4k images but just immediately compress them down using your trained PCA transformation. So, although you have less numbers than before, you still need to collect the original data. A real feature selection process would be able to do something like: "the proximity of the closest Applebees is not important to predict house prices, you should probably stop wasting your time calculating this number". As others have mentioned, L1 regression or some statistical procedure to identify useless features is typically how this is done. I would also add that domain knowledge is probably your #1 feature selection because we have to restrict the variables we input in the first place and which data we prioritize is inherently selecting the features.
- dkersten 5y agoSo would this be a more-or-less correct tl;dr: Dimensionality reduction is a compressing data in a way that retains the most important information for the task Feature selection is removing unimportant information (keeping/collecting, or selecting, only the important parts) Both cut down on the amount of data you end up with, but one does it by finding a representation that is smaller, the other does it by discarding unnecessary data (or, rather, telling you which data is necessary, so you can stop collecting the unnecessary data).
- naijaboiler 5y agoSo they are functionally doing the same thing, reducing the amount of data used. I find the debate about what we should call it useless and pointless. Understand what it is, what it doing, what it s limitations are, and use it appropriately based on your needs. Done.
- platz 5y agoif you read the actual article, dimensionality reduction by PCA does not retain the most important information about the data
- bllguo 5y agoindeed, predictable and disappointing how the discussion devolved into pedantry. should have been obvious what the author meant (plus it is clarified at the very beginning of the article). I'm not sure if this is a ML practitioner vs. statisticians thing or what
- hervature 5y agoYes, I would say this is entirely correct.
- chombier 5y ago> so you can stop collecting the unnecessary data I think that's the key, thanks. But still, if some inputs are redundant shouldn't this be somehow apparent in the eigen-vectors/values of the covariance matrix (making PCA an indirect feature selection algorithm)?