2 ms·
The result is in the large sample limit, which is pretty much never the case for high dimensional datasets like the ones most popular for ML these days (images,
by highd 9y ago
The result is in the large sample limit, which is pretty much never the case for high dimensional datasets like the ones most popular for ML these days (images, audio, text). It doesn't mean what the parent thinks it means.
- taeric 9y agoAh, so that means it doesn't mean what I also took it to mean. :) Know a good reading on this?
- apathy 9y agoYou are proposing that a reduced dimensional projection of a large dataset cannot approach this limit? I.e. expose the underlying low rank of nearly any huge sparse data matrix with an SVD or NMF. Enable fast recovery with a shitty (CS-wise) hash function. Recover most of the information about an observation's neighbors in a fraction of the time taken by many other approaches. What's popular for ML benchmarking these days is not necessarily the same as what's needed for a specific application. It's a useful proof to keep in mind before prematurely optimizing with overly complicated approaches.
- highd 9y agoYou are of course free to try simple dimensionality reduction and nearest neighbors, and if that works on your problem that's fantastic. To the research community, though, problems where approaches like that work were considered "solved" decades ago. And of course, in industry, if there's a chance of that working it's tried. But no one's building self driving cars with PCA and LSH.