5 ms·
As a researcher, I've even had difficulty making sure the code matches the algorithm as described in the paper. Getting old code running and validating that it'
by techwizrd 3y ago
As a researcher, I've even had difficulty making sure the code matches the algorithm as described in the paper. Getting old code running and validating that it's correct so you can compare is really non-trivial.
- bravura 3y agoGo to arXiV and download the latex source. Plug it into ChatGPT and ask it to implement in pytorch. I've really enjoyed this approach.
- Calamity 3y agohow well does this actually pan out for you?
- bravura 3y agoI've used it with great success.
- techwizrd 3y agoCan you provide an example of where this has been successful? I've spoken with many researchers and grad students only to find that there was a critical typo in the algorithm or undescribed setting (e.g., only converges when a learning rate scheduler is applied) in the code. I'll see the same algorithm implemented differently in different repositories. This is even the case for papers I've found with thousands of citations. It can be tremendously difficult to reproduce the results of papers, especially when they may require large amounts of compute that small researchers don't have. And I don't know whether it's a bug in my code or the paper or the algorithm.
- bravura 3y agoI mean, what you describe is an unfortunate and unavoidable issue in academia (and in the world in general). GPT4 doesn't work magic here, of course. You still have to: 1) Understand the work and the motivation. (GPT4 can help by playing the role of junior PhD if you can play the role of astute advisor.) 2) Sniff out things that are underspecified or seem wrong. (GPT4 also can help here, see above.) 3) Email the authors with questions, compare against shitty published codebases, etc. depending upon how gnarly/rushed the prose is. With that said, it's also "research smell" (compare "code smell") if a paper is so hastily written and undercited that you're the first person replicating it. And maybe instead of going for "my new bleeding edge approach that got 0.1% score better than boring old model with 50 cites", you probably should just implement boring old model. So, where this has been successful for me is in implementing denoising diffusion for different problem domains. Given that there is broad literature on denoising diffusion, when some things are underspecified you can start looking at best practices for other researchers. Alternately, other things like specific transformers etc. Basically what I'm saying is that if you are trying to reimplement something that is so niche, it's like catching butterflies. A better research agenda involves working within a particular field of study where there is supporting evidence and approaches to compare and constrast to. This goes without saying, regardless of whether an AI is involved or not.
- mistymountains 3y agoI’d be very careful about this.
- godelski 3y agoI have a habit of rewriting code because it is hard for me to understand without doing this and frankly, it isn't uncommon for the algorithms to be different. But the most common thing is that a small one-liner is a critical component to making the model work. Sometimes this is a one-liner in the paper, sometimes this is a one-liner in the paper that they built their work on, of the paper they built their work on, of the paper they built their work on. People building on top of one another's code definitely helps make projects get up and running faster, but I do wonder if it makes research faster and/or better. If a secret sauce gets lost in the game of telephone, it can be hard to know how much is the secret sauce. Especially when comparing to works that don't use that line. Sometimes the secret ingredient gets lost and similarly the gains. It's also not uncommon for this secret ingredient to not be known, or acknowledged. Especially if it is "well known" (won't be for long).
- techwizrd 3y agoThere is so much "secret sauce" that we're unconsciously relying upon. I learned a lot of secret sauces by reading through various implementations, "logbooks" [0] by various researchers, and so on. It definitely makes AI/ML feel more like an art than a science, and it makes it very difficult to re-implement in another language (e.g., Julia) or framework. 0: Such as https://github.com/facebookresearch/metaseq/blob/main/projects/OPT/chronicles/OPT175B_Logbook.pdf https://github.com/facebookresearch/metaseq/blob/main/projec...