4 ms·
> because things were seriously getting out of hand. Whatever that means.
by gaur 11y ago
> because things were seriously getting out of hand.
Whatever that means.
- clishem 11y agoI thought it meant a filter for nonsense articles, but apparently arXiv is not very searchable anymore.
- karpathy 11y agoImagine waking up every morning with 50 new arxiv papers uploaded that night. You panic and quickly scan through the papers - any of them could be very related to your research, or scoop your latest idea, or have good ideas you can use in your own work. Arxiv makes no attempts to filter these for you, so it's up to you to carefully scan through this unlabeled list of paper titles. You eventually find 3 papers that you have to read and put them on your list. You manage to read 1 that day. Next day you wake up and 50 new papers are up. You iterate for a few weeks and suddenly you have a toread list of 20 papers and 100 new arxiv papers just came in that evening. That's what's currently happening in research at least in deep learning (but I imagine more widely too), especially around big conference deadlines, and that's what I label "things seriously getting out of hand". That's a first use case. The second way things are out of hand is that you remember this paper from 3 years ago that was very related to this one, but can't remember it's name anymore. Here you can sort by similarity to any paper, and usually these papers come up on top of the sorted list. This is also useful for finding related work. Another use case is a peace of mind that you somehow did not miss some papers that you definitely should know about. Google Scholar is supposed to have similar features: it emails you papers it thinks you would be interested in and can in principle show similar papers. I don't know what they do internally but these features are quite terrible and low quality in my own experience compared to what I get here. More generally the amount of innovation in Google Scholar over the last few years is sadly either zero or negative (but overall I still get nightmares about what would happen to academia if Google pulled a Google Reader with Scholar). For arxiv-sanity it's tfidf vectors of bigrams from full text of each paper and I do L2 lookups for similarity ranking and train personalized SVMs for people for recommendations. The results are, at least for me, significantly better.
- jakub_h 11y agoWhat about using deep learning to properly classify arxiv papers about deep learning (and other things, perhaps)? ;)
- dalke 11y agoI have no real point here, only historical commentary. I've been reading papers from the 1960s, which is when the term "information explosion" was coined. People then were struggling to stay current with the literature, and thought 'things were seriously getting out of hand.' This was the start of abstracting services, like ISI, where you could even arrange the results of a keyword search of all the new papers to be sent to you each week - a clear predecessor to personalized RSS feeds. Going back even further to the immediate post-war era, the library systems of the time, which were structured around books and journals and organized by topic, couldn't keep up with the deluge of research reports which cut across multiple topics. The field of information retrieval, using first punched cards and then computers, started because the publication flow was 'seriously getting out of hand'. Or for a specific example, after high T_c superconductors were discovered in 1986, there was a mad rush of interest as solid state physicists from around the world explored the new territory. A Google Scholar search for "high temperature superconductor" finds: 1986 - 846 publications 1987 - 2 600 1988 - 3 900 1989 - 4 780 1990 - 4 870 1991 - 5 250 That's 14 papers per day, any one of which might be "very related to your research, or scoop your latest idea, or have good ideas you can use in your own work." Granted, 14 << 50, but that doesn't include some of the papers about "high Tc" which don't use the whole phrase. Also, those are 14 peer-reviewed papers per day, so there has been some filter, and experimental research in high Tc research requires more equipment than deep learning. Think of my comment as a reminder that things have been out of hand for most of a century, and dealing with that deluge emotionally connects you to the headache that generations of researchers before you have had to suffer with. :)