5 ms·
PCA is pretty close to optimal in situations where you want to preserve global structure and variance-covariance matrices are highly informative. This comes dow
by prionassembly 5y ago
PCA is pretty close to optimal in situations where you want to preserve global structure and variance-covariance matrices are highly informative. This comes down to the min-max theorem for Rayleigh quotients; I won't grandstand going over the math here. It won't work well for gene sequences or text embeddings because the desired structure there is local (UMAP seems to reign supreme these days for that).
An older explanation for PCA is that it's a sort of default factor model prior to rotations that introduce structural assumptions. (Varimax etc. factor analysis is really underrated in exploratory statistics; but then, by now data science training never introduces people to the FWL theorem, identification, etc. With large enough deep learning pretty much anything is possible -- also because using deep learning implies having truckloads of data -- but the median xgboost guy is way out of his element.)
- taylorius 5y ago"and variance-covariance matrices are highly informative" So - a distribution which is linear?
- kurthr 5y agoI don't think that the distribution needs to be linear... simply the correlation between N elements needs to be linear (or 1st order dominant). It works great for estimating multiple (e.g. different causes) small deviations (over multiple elements) from a global response. If you're looking at those larger eigen values, then you're looking at the dominant mode shapes of the eigen vectors of correlation. As others have said, it's better for estimation (not for classification).
- prionassembly 5y agoYou keep using that word.
- spekcular 5y agoA multivariate Gaussian, for example.