3 ms·Fair enough, but that still uses a lot more memory during training than what DeepSeek is doing.by eigenvalue 2y agoFair enough, but that still uses a lot more memory during training than what DeepSeek is doing.