4 ms·
There seems to be a big difference between efficiently training a "large-ish" model on 4-8 GPUs and a gigantic model on 1000+ GPUs. The same techniques might no
by bitL 4y ago
There seems to be a big difference between efficiently training a "large-ish" model on 4-8 GPUs and a gigantic model on 1000+ GPUs. The same techniques might not work due to different warm up steps, gradient overlaps etc.
All you can see running in the wild are quantized LLaMA variants (4 or 8-bit) whereas the original model is 32-bit.