3 ms·
The point about using FP32 for training is wrong. Mixed precision (FP16 multiplies, FP32 accumulates) has been use for years – the original paper came out in 20
by jwan584 2y ago
The point about using FP32 for training is wrong. Mixed precision (FP16 multiplies, FP32 accumulates) has been use for years – the original paper came out in 2017.
- eigenvalue 2y agoFair enough, but that still uses a lot more memory during training than what DeepSeek is doing.