3 ms·
(1022362 + 82432) gpu-hours / 2048gpus / 5 months ~= 15% uptime. That's only 0.08 nines of availability! I remember in one of their old guidebooks a lot of st
by SethTro 4y ago
(1022362 + 82432) gpu-hours / 2048gpus / 5 months ~= 15% uptime.
That's only 0.08 nines of availability!
I remember in one of their old guidebooks a lot of struggle to keep their 64 machine (512 gpu) cluster running this was probably 4x the machines and 4x the number of cluster dropouts.
- Tepix 4y agoThey may have thrown away some models that didn't turn out great.
- foobiekr 4y agoPoor GPU utilization even when available is the rule. Truly amazing. Staging of data is probably a huge part of it.
- mirker 4y agoIs it failures or is this some backfill/budget scheduling while everyone is sleeping?
- foobiekr 4y agoA lot of it appears to be non-streaming approaches to data distribution resulting in actual job behavior that looks a lot more like stage-process-clear batch jobs than what you'd want to hide the latency of data moves.
- pavelstoev 4y agoAt CentML, we profiled GPU utilization on a larger AI/ML research institute cluster. 10% to 45% range, mostly in 10% utilization range. We then offered them software optimizers (which do not affect model accuracy) to get to the 90% utilization for GPUs
- foobiekr 4y ago90% sustained utilization is quite amazing, and 10% is shockingly typical. I am a quite skeptical that this holds for training and very large data sets, of the sort where data placement comes into play, but if so, congratulations, and I hope things go well for you.