4 ms·
Yeah I thought I was reading the title wrong! For anyone that isn't aware, the role of pca is to create new (synthetic) features that represent the original fe
by rlayton2 5y ago
Yeah I thought I was reading the title wrong!
For anyone that isn't aware, the role of pca is to create new (synthetic) features that represent the original features.
It does not tell you which features of the original set are good for feature selection purposes.
- pas 5y agoso what's a good way/algorithm/strategy/method/technique to select features? (obviously asking for a friend who's not that familiar with this, whereas I've all the black belts in feature engineering!)
- z2210558 5y agoL1 regularisation is the usual way (see e.g. https://en.wikipedia.org/wiki/Lasso_(statistics) https://en.wikipedia.org/wiki/Lasso_(statistics))
- microtonal 5y agoThere are also some iterative methods like grafting (using feature gradients): https://www.jmlr.org/papers/volume3/perkins03a/perkins03a.pdf https://www.jmlr.org/papers/volume3/perkins03a/perkins03a.pd... and gain-based selection (using the improvement of the objective), see the appendix of: https://aclanthology.org/J96-1002.pdf https://aclanthology.org/J96-1002.pdf We used grafting for parser feature selection, for which it worked quite well: https://danieldk.eu/Research/Publications/ucnlg2011.pdf https://danieldk.eu/Research/Publications/ucnlg2011.pdf
- clove 5y agoPsychologists use factor analysis.
- ylks 5y agoFeature selection ought to be model-specific. Because a feature wasn't selected by Lasso (in a linear model) does not mean it cannot be useful in a non-linear model.
- tomrod 5y agoI work with ML regularly, and there is always something new to learn! Another commenter mentioned L1 regularization, which is useful for linear regression. You wouldn't use it for all classes of problems. L1 regularization has to do with minimizing error of absolute values, instead of squared errors or similar. I skimmed this article and thinks it's accessible: https://www.kdnuggets.com/2021/06/feature-selection-overview.html https://www.kdnuggets.com/2021/06/feature-selection-overview... PCA is a form of dimensionality reduction, but it doesn't select features for you.
- georgefox 5y agoI'm probably nitpicking your language, but L1 regularization is precisely that: regularization. (See https://en.wikipedia.org/wiki/Regularization_(mathematics)#Regularizers_for_sparsity https://en.wikipedia.org/wiki/Regularization_(mathematics)#R....) In your typical linear regression setting, it does not replace the squared error loss but rather augments it. In regularized linear regression, for example, your loss function becomes a weighted sum of the usual squared error loss (aiming to minimize residuals/maximize model fit) and the norm of the vector of estimated coefficients (aiming to minimize model complexity).
- tomrod 5y agoHey, I appreciate your correction! I wrote my comment late and night and definitely mashed the details. Your nice "nitpick" is a much needed correction to my inaccuracy.
- r-zip 5y ago> L1 regularization has to do with minimizing error of absolute values Not quite. It has to do with minimizing the sum of absolute values of the coefficients, not the error. The squared error is still the "fidelity" term in the cost function.
- newrotik 5y agocheck scikit-learn feature selection utilities https://scikit-learn.org/stable/modules/feature_selection.html https://scikit-learn.org/stable/modules/feature_selection.ht... besides techniques mentioned in other posts (l1 regularization) sequential feature selection (backward and forward) is quite common
- disgruntledphd2 5y ago> equential feature selection (backward and forward) is quite common This is fine if you have test/validation sets, but never, ever report p-values on the result of such a selection process, as they are incredibly biased.
- naveen99 5y agoThis is the basic promise of deep learning: resnet, unet, transformers etc… you replace feature engineering with deep learning models… otherwise feature engineering comes from domain expertise/ human experience. or trial and error from known features that work elsewhere.
- disgruntledphd2 5y agolasso is probably the easiest way to do it relatively quickly. Lasso is also known as L1 regularisation, and it tends to set the coefficients to a bunch of features to zero, hence performing feature selection. Note that if two predictors are very correlated, lasso may pick one mostly at random. Obviously one should do CV and bootstrapping to ensure that the results are relatively stable. In general though, there's no real substitute for domain expertise when it comes to selecting good features. edit: lasso is L1, not L2
- ylks 5y agoCheck out these posts: https://blog.kxy.ai/tag/feature-selection/ https://blog.kxy.ai/tag/feature-selection/. This one in particular compares a few methods on 38 datasets and has some Python code: https://blog.kxy.ai/adding-feature-selection-to-any-model-in-python/ https://blog.kxy.ai/adding-feature-selection-to-any-model-in.... Disclaimer: I wrote the original blog post.
- platz 5y agoIsn't the article in fact saying that the new synthetic features are not good features for training?
- galangalalgol 5y agoThe eigen values do give information as to energy, so if your selection criteria is simply to pick the linear component with the most energy, you can use PCA to select that feature and extract it with the corresponding eigen vector. The MUSIC algorithm is a classic example. Edit: having now actually read the article, the case I mention falls into the author's test of (do linear combinations of my features make sense and have as much of a relationship to the target as the features themselves).