4 ms·
A lot of the AI progress is stalled because these chips do not handle thermal contraction/expansion well at all, and end up permanently destroying these chips.
by vkaku 25d ago
A lot of the AI progress is stalled because these chips do not handle thermal contraction/expansion well at all, and end up permanently destroying these chips. HBM has been a failure at this scale with chips getting destroyed every 4 months or so during operations.
I'd like to see some actual science saying, here was the problem, here's how we solved it, here's the AFR data, here's this running after X cycles etc. Nobody has done this reliably yet. That entire industry is hiding the bodies.
- _joel 25d agoHow many contraction/expansion cylces do you normally see in a DC setting, generally? Is there a measure for that to baseline against?
- bearjaws 25d agoIIRC it's thousands per day on these systems, mainly due to high power density and low mass, even small lapses in computation (100-500ms) rapidly change the temperatures of stacked die. So even a GPU averaging 98% utilization may have thousands of cycles per day. Compared to a regular server blade it may be dozens or barely any at all.
- venussnatch 25d agoNaive question, but could this not be fixed by a scheduler? If it's only idle 2% of the time, give it busy work for that 2%
- adgjlsfhk1 25d agothe easier answer would be to remove the ability to idle