4 ms·
This is cool, but I've always wanted to know the reverse: how many arXiv papers that have an "algorithm" in them have made it into the codebase at a for-profit
by asdf_snar 5y ago
This is cool, but I've always wanted to know the reverse: how many arXiv papers that have an "algorithm" in them have made it into the codebase at a for-profit company? I estimate it at less than 0.1%.
Of course, no company would let you scan their codebase, but perhaps there is another way of estimating this number. I'm even willing to put out a $ bounty for a clear formulation (e.g. what constitutes an "algorithm"?) and answer to this question. I'm not sure if that sounds stupid (please raise your hand if it does!).
- krageon 5y agoMost papers don't make it into a codebase at a for-profit company because most algorithmic papers are fiction. The algorithms described can't or won't work, and in (almost) every paper I've read so far don't work (that is to say, if they can work, that means they don't work as described). I've seen one or two exceptions in a set of hundreds of papers related to classification in some way. A good rule of thumb I've found is that if a paper seems trivial (i.e. after the intro the solution seems obvious, as all good solutions do once you know them) it will probably yield a usable algorithm. So my estimate would be 0.01% or less. Edit: for clarity, I'm saying this because I think the question is not so interesting. This area of academia (and I suspect others as well, where I do not have the knowledge or tools to replicate them) produces largely trash.
- foldr 5y agoAlthough I think you might be exaggerating just a tad, it is amazing how often practical implementations of even "standard" algorithms aren't fully described in the original publication. For example, the original description of the Bentley–Ottmann algorithm† simply handwaves away a number of crucial edge cases. †An algorithm for finding the intersections in a set of line segments.
- krageon 5y ago> might be exaggerating just a tad Over the course of my last job, I've read over (but not far over) 200 scientific papers. Of those, I've rejected about three quarters because even the algorithm's description was enough to inform a reader that it would not work (assuming the reader understands the content well). The remaining quarter I've implemented in code and run, and of those runs I can count on a single hand how many algorithms even came close to doing what they promised. Writing this down, you are right that this is not cause to dismiss most papers as "probably fiction". It's uncharitable and a better reflection of how salty I am about the amount of work that went into sorting through and implementing all of that (mostly to absolutely zero results) than anything else :) > simply handwaves away a number of crucial edge cases That does also happen, but without those edge cases you can presumably still see the algorithm working as it should given the right inputs. From there it is a matter of refinement to come to an implementation that is okay - this I would count as a success.
- asdf_snar 5y ago> I've rejected about three quarters because even the algorithm's description was enough to inform a reader that it would not work > The remaining quarter I've implemented in code and run, and of those runs I can count on a single hand how many algorithms even came close to doing what they promised. I'm curious what you mean by "not work" here. Presumably such papers use examples to illustrate their algorithm. Were results not even reproducible on the authors' own (cherry-picked) examples? Or perhaps do you mean you threw a harder problem at the algorithm that cleanly fell within the set of problems the authors purported to address? I think the distinction is important, because the first case means the paper is just flat-out wrong. The second case is what causes all the trouble, because it's hard to convey to an academic that their algorithm, though in principle "correct", does not address the often-vaguely-stated problem in the introduction ("this algorithm has applications in X, Y, Z and related fields..."). They can just say "I'm advancing the field" or "it provides insight that could one day be more useful".
- asdf_snar 5y agoThank you for this reply. My experience aligns with yours, and I find your conclusion (the question is not interesting) insightful. You're probably right. I am an ex-academic that wrote useless algorithms. The implementations were mathematically correct and the algorithms did what we said they did, but the premise of the algorithm itself was flawed. To the extent that the algorithms were correct, the target problems were useless, and to the extent that the target problem was a "real-world" problem, there were a million things unrelated to the mathematics that made the contribution inconsequential. The reason I wondered about such a study is that I know many academics who wouldn't believe the situation as dire as you describe it (though I very much agree it is <0.01% -- I just wanted to be conservative).