3 ms·
I have no knowledge of such things, but it seems they run Cuda jobs on about 150 nodes? But why do they have so many problems to keep this cluster stable? Netw
by zubspace 4y ago
I have no knowledge of such things, but it seems they run Cuda jobs on about 150 nodes?
But why do they have so many problems to keep this cluster stable? Network failures? Bad GPU's? Bad drivers? Bad software?
Running fixmycloud and going after all those cryptic errors every day seems like a nightmare to me...
- semi-extrinsic 4y agoSeems kind of par for the course for an HPC cluster, no? It makes sense to think of these things like Formula 1 cars, they are trying to eke out the absolute maximum performance, and reliability suffers because of that. "Ordinary" cloud is more like a Toyota where you optimize for fuel economy and low maintenance.