3 ms·
A Google Nature Paper has not been replicated for over 3 years, but I'm the one fabricating stuff :D Making a novel claim implies its *_claimed_ replicability.
by twothreeone 2y ago
A Google Nature Paper has not been replicated for over 3 years, but I'm the one fabricating stuff :D
Making a novel claim implies its *_claimed_ replicability.
"You did not follow the steps" is calling them idiots.
The only inference I made is that he's pressed to comment. He could have said nothing.. instead he's lashing out publicly, because other people were unable to replicate it. If there's no problem replicating the work, why hasn't that happend? Any other author would be worried if a publication about their work were saying "it's not replicable" and trying their best to help replicate it.. but somehow that doesn't apply to him.
- griomnib 2y ago“You can only validate my results if you have an entire Google data center worth of compute available. Since you don’t, you can’t question us.”
- jeffbee 2y agoWe're actually talking about the difference between Cheng using 8 GPUs and 2 CPUs while Google used 16 GPUs and 40 CPUs. These are under-your-desk levels of resources. Cheng et al authors are all affiliated with UCSD which owns the Expanse supercomputer which is orders of magnitude larger than what you would need to reproduce the original work. Cheng et al does not explain why they used fewer resources.
- griomnib 2y agoThat’s a fair complaint then.
- phonon 2y agoNo it's not. They ran it longer instead.
- jeffbee 2y agoThe 2022 paper pretty explicitly says that runtime is not a substitute. They say their best result "can only be achieved in our 8-GPU setup".
- phonon 2y agoI assume you mean Fig. 6 here?[0] But that was explicitly limited to 8 hours for all setups. Do they have another paper that shows that you can't increase the number of hours of a smaller GPU setup to compensate? [0]https://dl.acm.org/doi/pdf/10.1145/3505170.3511478 https://dl.acm.org/doi/pdf/10.1145/3505170.3511478
- wholehog 2y agoThey also changed the ratio of RL experience collectors to GPU workers (~1/20th the RL experience collectors, 1/2 the GPUs). I don't know what impact that has --- maybe each GPU episode has less experience? Maybe that makes for an effectively small batch size and therefore more chaotic training? But either way, why change things when you can just match them exactly?