3 ms·
One thing which works against the "cite everything" approach is that most of the major conferences have page limits of 8-10 pages with 1 page bonus for referenc
by kastnerkyle 11y ago
One thing which works against the "cite everything" approach is that most of the major conferences have page limits of 8-10 pages with 1 page bonus for references. That means if you go over 1 page of references for (at least NIPS) then you cut into the meat of the paper, reviewers look on in disdain and give poor marks, etc. So you have to actively prune for the most recent and directly relevant citations many times, which sometimes counts out semi-relevant but older work in favor of more relevant recent work.
Much of Dr. Schmidhuber's work is very interesting and especially relevant now that RNNs are really heating up again - but it is sometimes hard to figure out exactly which of his papers to cite because many are partially relevant. And having a full page of only Schmidhuber citations is no good either...
Speaking as a member of the Montreal lab, I am much more up to date with the work that happens here - so it is hard to fight the natural tendency to cite recent papers you know (since they all came from work you know of, cause you were there). Notice too that all 3 (Hinton, LeCun, and Bengio) worked directly together at some point, and collaborated often beyond that. So a version of this is in effect, whereas Juergen has been more separated (both geographically, and work focus wise) than the other 3. NYU Toronto and Montreal are all in an 8 hour triangle!
Not to take anything away from his points (I try to cite as many of his papers as possible without seeming ridiculous, generally) but these are the general factors at play. We cannot possibly cite every paper in the field, and shining the light on new works can be more important than citing older work AS LONG AS there is no claiming as a pure innovation work that was already done "in the nineties".
Claiming to improve some technique or take it from curious to usable is more than fine - but given the recent deep learning hype even recent papers are getting overshadowed by others claiming some new innovation which already exists in very current literature.
Especially given the work that is coming out of industrial labs (Google, FB, MSR, etc.) it is fairly frequent to see the same model being touted as new (with minor citations if lucky) when the exact same technique first appeared 6 months ago. Being well-read is not an option as an academic - it is a requirement! The PR machine of these companies is unfortunately very effective at dominating the airwaves if you have competing or related work, especially if you are not from a school with good press e.g. MIT, Stanford.