4 ms·
I took a stab at trying to interpret the topics output by this run of LDA. Green is one the clearest: generally convolutional deep nets, image classification, e
by zackchase 12y ago
I took a stab at trying to interpret the topics output by this run of LDA. Green is one the clearest: generally convolutional deep nets, image classification, empirical work.
Brown seems to have picked up on linear algebra. "Vector", "matrix", "tensor" and "decomposition" all get consistently labeled brown, as do "eigenvalues", "orthogonal" and "sparse".
The rest are not as useful. Black almost always has "number", "set", "tree" and "random", but little else. Purple at times seems to signify topic modeling, but also contains "neural" and "feedforward". Blue seems to be the stats topic, containing "Bayes", "regression", "gaussian", and markov processes. But it also contains random words like "university" and "international".
Overall, very interesting. I wonder if these topics would be even better defined with a higher setting of k.
- dustintran 12y agoYup, it seems k was fixed since the first time these scripts were made for NIPS 2012 (?). Some of the more well-established advances since LDA would also likely help, like HDP.
- taneliv 12y agoKarpathy had a different interpretation (in the green bar at the top of the page). For example, purple would be neuroscience. In addition to adjusting k, another change that might be interesting would be to include also previous years' papers in the model estimation. Changes in component (topic) weights year-over-year could perhaps reveal something about the topics, or the papers.