2 ms·
Yep. Even if the initial seed for parameter init, the example shuffling seed, etc were constant, the distributed nature of training (and syncing the model acros
by Scene_Cast2 3y ago
Yep. Even if the initial seed for parameter init, the example shuffling seed, etc were constant, the distributed nature of training (and syncing the model across machines) would kill reproducibility. Not to mention resuming from checkpoints after gradient explosions, etc.
- monocasa 3y agoI've heard from ML engineers at larger shops that reproducibility is key to working at scale. That's how you track down "this training regime went to shit because of something we changed" versus "this training regime went to shit because on of the GPUs training it is starting to fail".