3 ms·
Deepseek was trained on Nvidia GPUs H800. Nvidia is still selling GPUs to China and the reason even the reduced chip performance is easily negated by scaling.
by Jlagreen 2y ago
Deepseek was trained on Nvidia GPUs H800. Nvidia is still selling GPUs to China and the reason even the reduced chip performance is easily negated by scaling.
US government is stupid because they asked for certain limits on chip base first but AI uses GPU clusters. In a GPU cluster you don't have full utilization anyway so slower GPUs don't matter as much as slower networking. China still gets pretty high bandwith Nvidia HW for building large clusters for training/inferencing.
Chips from Huawei still seem to be way too unstable for the job. In training/inference stability is even more important than performance. Imagine you have a fast chip but it can't run without errors for 2 months and you training never gets done. That's DoA.