12 ms·
Deep Learning Papers Are Kinda Bullsh-T
- dunefox 4y agoThere really is a lot of cherry picking, etc. going on in this area. Papers released without code and weights or even data make reproduction and validation nearly impossible.
- forgingahead 4y agoYeah it's stunning to me that people can apparently run experiments with code, produce results with code, write a paper about that code, and then release poorly written prose in a garbled way without also releasing that code (or at the very least releasing a video demonstrating results).
- kfarr 4y ago...And get a PHD for it!
- the_only_law 4y agoI was under the impression this is just how academia worked nowadays.
- harha 4y agoThere are really just two situations if the solution generalizes well - and if it doesn’t it might be worth just mentioning that and to move on: 1. similar open data exists, great, just publish a sample implementation 2. if not the first task is to generate such an open data set Edit: formatting
- deleted 4y ago[deleted]
- time_to_smile 4y agoThere's also the problem that most complex neural networks are highly sensitive to initial weights. My friends and I have frequently tried to reproduce famous papers and it's remarkable how often getting the initial settings nearly exactly correct is the key to achieving the targeted bench mark. This is a problem because cherry picking is essentially built into the frame work. If I was building ranking algorithm and just kept picking a random seed to arbitrarily sort a list of numbers until it was correct, most people would consider that obviously cheating. However if I did the same thing but stuck 3 dense matrices between the seed and the list to be ranked it would considered AI.
- kache_ 4y agoNo demo, no read, no cite
- biesnecker 4y agoAm I the only person who read that title as `Bullsh[T]`?
- paskozdilar 4y agoIf I ever write my own shell I'm gonna call it `Bullsh`.
- f137 4y agoCollecting the data is costly. I don't believe we will see many private datasets openly released.
- ssivark 4y agoIMHO, it’s very important for papers to lay out the idea lineage and contextualize their incremental progress (analogous to a codebase commit history). For whatever reason, ML seems to prefer the practice of framing each paper as a shiny nugget uniquely disconnected from the ecosystem of ideas, and with is own fancy name (as if it was born a fully developed rockstar). Imagine isolated code dumps without the shared history leading to nightmare merge conflicts… I think this makes its really hard for anyone not steeped in experience to parse through the outputs of the spraying firehouse, and organize their thinking rigorously — thereby fragilizing the field’s intellectual output in a vicious cycle.
- bravura 4y agoYeah, and don't just dump shitty impossible to run research code on github with a half-assed README. Give me a one liner huggingface or torchhub, or a working google colab. Or I'm probably nexting your work and trying the second and third best model instead.
- auxym 4y agoCan you even control which version of python + external dependencies you get in a colab? Whatever you publish and works now will not necessarily work in a year or two (which is often how much time it takes to get a journal paper published).
- kordlessagain 4y agoThe Yolo7 Github has a good amount of spelling and capitalization errors. I guess they only looked once.
- madduci 4y agoWhy are you blaming an implementation project here ? YOLOv7 is based on multiple papers?
- lesquivemeau 4y agoYou seemed to have missed a joke here. I guess you only looked once.
- mmmmpancakes 4y agoone possible explanation: if you open source all your data and code you likely reveal the secrets that aren't in your paper that give you a competitive research edge. Writing papers so that you are telling enough to get a good publication while obfuscating the details that give you an edge in your research fiefdom is a bit of a dark art in academia. I think it should be discussed more.
- softwaredoug 4y agoMy trust in most academic papers is very low. Exceptions are made for certain well known groups and authors, but generally I’m not going slog through difficult to understand paper for results that can’t be reproduced - not even to mention recreated in a production setting.
- mjburgess 4y agoThere does need to be some kind of reckoning here, pseudoscience is ballooning and dwarfing the actual practice of science. We're long past the point where researchers "should know better" --- we're now into Nature publishing this BS.
- TrackerFF 4y agoFirst hurdle is to simply get the (more often than not, Python) dependencies to work. I've worked on reproducing some relatively simple DL programs - written by academics I know - where I've literally spent days to weeks just to get all the dependencies right. And I've had direct contact with them - which may absolutely not be the case for other people. I don't know why DL libraries are so afflicted by this, maybe things just move so fast. But it is such a pain in the ass.
- SQueeeeeL 4y agoYou hit the nail on the head. DL library maintainers basically have no respect for backwards compatibility and ensuring everything works. New versions are pushed out on a weekly basis that break existing APIs, and no one really cares because dependency management has been abstracted so far away maintainers don't even understand the repercussions to this 'move fast, break things' mindset (namely, lots of broken software)
- iakov 4y agoIMO it's not DL libraries, it's Python. Python sucks at managing dependencies. It's a hilarious mess of pipenv, prose, conda, vex, pex, shmex and god knows what else is hot now. It seems that every time I want to write a simple Python utility, there is a new way to install and track dependencies.
- pbourke 4y agoOnce you learn the basics of pip/venv that should mostly work for everything. Make a new venv for everything and don’t pollute the global environment and it should be fine.
- mbrudd 4y ago> Make a new venv for everything and don’t pollute the global environment and it should be fine. This just proves the point that Python sucks at managing dependencies, which exacerbates -- perhaps even encourages -- the reproducibility issues being discussed.
- davidktr 4y ago>Science, since time immemorial, has relied on the systemic replication of any presented result or finding. Reproducing experiments and their reported results remains a cornerstone of the validation of any scientific theory. No, no it really hasn't. It has relied on the ability to make predictions based on pusblished theories, methods, laws etc. Even for hard-science experiments it's not even clear how you could record all the required knowledge to replicate an experiment. Every configuration, every machine, every particle in the air, every bit of software. I really wish people engaged more with the actual history of science instead of what they believe it to be. edit. To give a little more meat to my rant here's a good reading (https://plato.stanford.edu/entries/scientific-reproducibility/#ReplDistFeatScie https://plato.stanford.edu/entries/scientific-reproducibilit...): "If replication played such an essential or distinguishing role in science, we might expect it to be a prominent theme in the history of science. Steinle (2016) (...) claims that the role and value of replication in experimental replication is 'much more complex than easy textbook accounts make us believe' (2016: 60), particularly since each scientific inquiry is always tied to a variety of contextual considerations that can affect the importance of replication. Such considerations include the relationship between experimental results and the background of accepted theory at the time, the practical and resource constraints on pursuing replication and the perceived credibility of the researchers. These contextual factors, he claims, mean that replication was a key or even overriding determinant of acceptance of research claims in some cases, but not in others." The history of replications is extremely nuanced. Empirical results and by extension replications are one line of argument in scientific discourse, but by no means the only one. I personally hold that valid predictions in the context of interesting problems are where it's really at. In the "Structure of Scientific Revolutions", Kuhn argues that at some point paradigms cannot make THESE kind of predictions anymore. Revolutions do not happen because of failed or missing replications. Therefore, stating science "has relied on replication" is historically and epistemologically false. It's also misleading because the replication crisis happens due to a lack of theory and misguided incentives, not because some discipline has left the holy path of finding truth.
- unixhero 4y agoHe refers to scientific positivism. https://plato.stanford.edu/entries/logical-empiricism/ https://plato.stanford.edu/entries/logical-empiricism/
- Kalanos 4y agohuh? most of this article gripes about the intricacies of publishing a paper. use https://docs.aiqc.io https://docs.aiqc.io for reproducible protocols
- twak 4y agoAcademics are judged by the publications not their implementations, so the system favours over-sold manuscripts and it-ran-once implementations. Until funding is conditional-on (and provided for) robust well maintained code it will remain challenging to get reproducibility. Frequently the PIs (bosses) will not even glance at the repositories written by junior members, probably can't read code anyway, and certainly won't allocate time for their maintenance. Even worse, most academics who do publish code have never been exposed to real world software engineers, their techniques, or tools.
- InefficientRed 4y agoThe basic issue is the labor. We don't pay for good science. Suppose I told you to develop good software that's novel enough to publish about, but only gave you enough budget to pay your SWEs a maximum of $30K/yr. That's one zero, for those reading quickly. Additional non-beneifts: 1. Unlike literally every other job in the country, you don't have budget to pay FICA taxes for your employees, and tax code allows this. This means your employees don't even have the USA's paltry social safety net to fall back on if they are hit by a bus or graduate into a massive recession, and their years working for you do not count toward social security or medicare retirement benefits. 2. Obviously, there is no budget for 401K retirement benefits 3. No CoL raises 4. Healthcare benefits will be paltry. 5. Your SWEs need to serve as a teaching assistant every once in a while. This likely means grading homework and a few late evenings of grading exams. No overtime for those late nights, obviously. 6. All travel, which is mandatory and often international, must be paid by the employee up front and reimbursement can take 1-3 months. We don't trust $30K/yr drones with corporate cards. Good luck making rent after a conference :) Just to reiterate: You need to hire SWEs. You pay $30K/yr (less than some Amazon warehouses!), benefits package is literally worse than a part-time gig at a supermarket or fast food joint, and your employee is expected to give you $2K-$4K loans a few times a year while living paycheck to paycheck. I just roll my eyes hard when I see complaints about garbage research code. Almost everyone in my PhD cohort had FAANG or finance offers; we were all taking 5x-10x paycuts to work on interesting problems and do science. If you want productizable research prototypes, hire PhDs to do science for you. (And I say this, for the record, as a rare PhD who during their phd wrote code that is well-documented, well-maintained, and still used by dozens of companies for business-critical processes many years later.)
- amitport 4y agoThis post is kinda Bullsh-T. Basically all claims are unfounded, prior work is missing, and it is filled with filler content (i.e., boilerplate? I'm not sure how to call it) instead of providing value.
- Imnimo 4y ago>However, this might not be the case. Let’s take for example the Fall 2021 Reproducibility Challenge - an event designed to encourage reproducing recently published research from top conferences. Only 43/102 (~%42.16!) of the papers entered into the double-blind peer-review process were accepted – which means that more than half the papers, despite being written with reproduction as a priority, couldn’t actually be reproduced. I don't think this is what a rejection means. Papers are accepted and rejected from the challenge depending on whether they do a good and thorough job of attempting to replicate the original work, not depending on whether or not they succeed.
- dang 4y agoYou can't do that with titles here. Please see the site guidelines.
- lpasselin 4y agoOne detail some might not realize is the fact that research code is often a heaping pile of garbage written by a single graduate student. Some are ashamed of their code and simply don't want to publish shit code. Also, strategically, it is probably better to _not_ publish code than risk being rejected by a future job interviewer because your research code is shitty and you didn't prioritize refactoring it. With that being said, this is not an excuse for refusing to share paper code or making sure the experiments are reproducible.
- twak 4y agoOther causes include the pressure to publish quickly in ML (while your approach is en vogue), with small teams, before your funding runs out, while hitting conference deadlines. In these situations, I have suggested releasing anonymous implementations after the paper is accepted just to get the code out there. I am not certain this is the right thing to do!
- n_time 4y agoA large portion of this is due to the corporate enclosure of deep learning and machine learning that has occurred over the past 10 years. This, combined with the scale required for a lot of deep learning research, means that neither the code nor the resources required for reproduction are accessible outside of corporate labs.
- lkois 4y agoMy last job was centered around trying to reproduce the results of deep learning academic papers, and maybe porting them to other frameworks or platforms. Even WITH code supplied by the authors, this was always a struggle. It'd usually take about a week or so just to get their github project out of dependency hell and actually run at all. And if it needed to be reproduced in another framework, I'd really really want some kind of demo code just to clarify what exactly the authors were trying to describe. Especially if their descriptions had holes or discrepancies that only became clear when trying to fit the pieces together. I remember trying to reproduce a couple of object tracking papers from the same authors, one with an overly complex and poorly defined feature set, the other with a glaring mistake/omission that forced my team to redesign the model because they described using a certain layer type in a way that made no sense. There were a few good exceptions that provided nice code, but difficult reproduction seemed to be the norm.
- clircle 4y agoEven if deep learning papers were more reproducible, there's still little confidence that the flashy new technique will work well for _my problem_. I'd rather see machine learners work on techniques that work robustly than one that are 1% better on a very narrow problem.
- axg11 4y agoDeep learning is one of the most reproducible areas of science. That might seem like an insane statement until you spend time in a biology (wet) lab. Experiment protocols are often poorly documented and access to materials needed to reproduce experiments are highly uneven. There is even more cherry picking when it comes to biology, especially since biologists don't often have the statistical knowledge to know better. I'm not writing this to defend deep learning. Reproducibility is an incentives problem across ALL of science. We value novelty and prestigious publications over everything. Nobody wants to fund "boring" research that reproduces existing results. To fix academic science, we need to reward reproducible research and fund groups such that they're capable of performing it.