3 ms·
Transmission speeds aren't fast enough for this, unless you crank up the batch size ridiculously high.
by ImprobableTruth 3y ago
Transmission speeds aren't fast enough for this, unless you crank up the batch size ridiculously high.
- FeepingCreature 3y agoLoRA training/merging basically is "crank up the batch size ridiculously high" in a nutshell, right? What actually breaks when you do that?
- brrrrrm 3y agoCranking up the batch size kills convergence.
- FeepingCreature 3y agoWonder if that can be avoided by modifying the training approach. Ideas offhand: group by topic, train a subset of weights per node; figure out which layers have the most divergence and reduce lr on those only.
- brrrrrm 3y agoA provable way to recover convergence is to calculate the hessian. It’s computationally expensive but there are approximation methods.