7 ms·
Like some of the other ML/AI posts that made it to the top page today, this research too does not give any clear way to reproduce the results. I looked through
by gigantum 7y ago
Like some of the other ML/AI posts that made it to the top page today, this research too does not give any clear way to reproduce the results. I looked through the pre-print page as well as the full manuscript itself.
Without reproducibility and transparency in the code and data, the impact of this research is ultimately limited. No one else can recreate, iterate, and refine the results, nor can anyone rigorously evaluate the methodology used (besides giving a guess after reading a manuscript).
The year is 2019, many are finally realizing it's time to back up your results with code, data, and some kind of specification of the computing environment you're using. Science is about sharing your work for others in the research community to build upon. Leave the manuscript for the pretty formality.
- Havoc 7y ago>any clear way to reproduce the results. Given that it's evolved I'd imagine this is a given? Or more accurately you could probably duplicate some kind of emergent behaviour but it would be different given different randomized parameters
- lysium 7y agoUsually you use an RNG for which you can publish the seed. So, although it’s random, you can reproduce the results.
- tastroder 7y agoGlancing through the paper it seems like they use the recent Transformer model. Does whatever underlying stack they use expose something to share RNG seeds and the exact hardware optimizations your environment applies during training? Otherwise "publishing the seed" sounds nice but might not be as trivial as the phrase suggests.
- arthur_pryor 7y agoreproducibility should be something that's baked into an experiment's design. so, if their experiment was designed such that reproduction is inherently difficult, they should have designed it in a better way, and they should've used a toolset that wouldn't run into that problem. a non-reproducible experiment isn't necessarily completely without value, but it's a thing that everyone should look askance at till it proves its worth. (apologies if my comments don't apply to this experiment and if it is reproducible -- i didn't have time to read through the OP, but i thought this reply was still a worthwhile response to its specific parent comment)
- tastroder 7y agoNo that's absolutely a fair and true point, my comment was more pointed at the RNG aspect. I have not looked into this specific one either but normally people would hopefully not publish their best randomly achieved run if the system cannot reproduce it or similar results. That being said the paper in question doesn't seem to reference open source code anyway so I guess my point was kind of moot, apologies.
- lysium 7y agoYou want to be able to set the seed if only you want to be able to debug your program. Pseudo random is sufficient for these models and is independent of any hardware settings. You should not share your random source between concurrent threads, though, but that’s good practice anyway.
- nl 7y agoFor the most part, yes. There are specific CUDA operations which are not guaranteed to be reproducible though, as well as some CuDNN operations which are non-determanistic without performance sacrifice, and this does cause real problems. See https://pytorch.org/docs/stable/notes/randomness.html https://pytorch.org/docs/stable/notes/randomness.html for some reasonable docs on this.
- seppel 7y agoThere are many CS conferences where you can/should submit a VM image to reproduce the results. See, e.g.: http://cavconference.org/2018/artifact-submission-and-evaluation/ http://cavconference.org/2018/artifact-submission-and-evalua...
- londons_explore 7y agoMost machine learning accelerators have a few non-deterministic operations. The chances that you could run trillions of floating point operations through a GPU and get a bit-for-bit identical result is low.
- tempguy9999 7y agoReally? I'm not an ML guy so in simple terms, what are these non-deterministic ops? Or are you saying GPUs can be expected to be, basically, faulty?
- londons_explore 7y agoBoth. Some operations split and join data in non-deterministic ways (especially the order of operations, leading to different floating point rounding). If you shard across multiple machines, weight accumulation order will depend on network latency for example. Also, GPU's aren't anywhere near as reliable as CPU's when it comes to being able to run for hours without any random bit flips/errors.
- tempguy9999 7y ago> ...split and join data in non-deterministic ways ... to different floating point rounding Ah, of course! A very timely reminder, thanks! > GPU's aren't anywhere near as reliable as CPU's when it comes to being able to run for hours without any random bit flips/errors. Now that's worrying. A bit flip can't be expected to be skewed towards any particular bit within a float, so it could easily happen in the exponent, skewing a single value by orders of magnitude one way or the other. Combine that with the rest of your 'good' results and yuck. That's very concerning. Thanks for the warning.
- londons_explore 7y ago> A bit flip can't be expected to be skewed towards any particular bit within a float, Actually, I think they are - for example, the exponent path through an adder/multiplier is typically shorter, so when operated close to clock speed limits, the exponent is more likley to be correct. (I've not actually verified the above on real hardware)
- vanattab 7y agoIs it not possible to use the same seed and random number generator to reproduce the results accurately?
- visarga 7y agoYou got to be careful. RNG is being used to initialise the layers but also for mini-batch selection. They are usually different RNG's.
- threwawasy1228 7y agoMore of what the point is I think is that they don't go into any meta-analysis of big changes that were seen in many of the trials. They don't try to isolate specific mechanisms that formed in a majority of trials that almost made it to this stage for example. They just don't really go into any analysis of the failure trees in trial dataset at all. IMHO this is probably just a case of them trying to stretch this out across a bunch of different papers, and this is just the announce paper. Which is a shitty practice, but the current academic environment encourages taking good findings and puffing them up into multiple incomplete papers rather than one well-done paper.
- deleted 7y ago[deleted]