4 ms·
Possible and has been done, but super-slow and inefficient resulting in long training times for small models. To keep compute occupied you need to pass gradient
by coolspot 1y ago
Possible and has been done, but super-slow and inefficient resulting in long training times for small models.
To keep compute occupied you need to pass gradients very fast.
- pk-protect-ai 1y agoDo you mean this one? https://blog.lambdaclass.com/introducing-demo-decoupled-momentum-optimization-for-efficient-distributed-llm-training/ https://blog.lambdaclass.com/introducing-demo-decoupled-mome...
- reactordev 1y agoThis is what piqued my interest in the first place
- reactordev 1y agoYes but could you break it up into chunks of sets of gradients to compute? I know that compute needs the full chunk to compute a set. Again, things I’m exploring but ultimately no different than just having the full dataset on disk and just scaling out compute nodes in ro mode.