3 ms·
All included it costs under 70$ for the 13B model. Training 65B now so will report what that will cost.
by sahil_chaudhary 4y ago
All included it costs under 70$ for the 13B model. Training 65B now so will report what that will cost.
- sillysaurusx 4y agoPlease do! Also please include how you’re calculating the costs.
- alex_sf 4y agoFor the 65B fine tune, did you add another A100 node? Or just drop batch size? Any chance you’re up to sharing the training parameters?
- sahil_chaudhary 4y agoDropping the batch size