3 ms·
What compute environments are used for LLMs like ChatGPT? Do you see this recent frenzy driving demand for HPC engineers?
by morelandjs 4y ago
What compute environments are used for LLMs like ChatGPT? Do you see this recent frenzy driving demand for HPC engineers?
- hackandthink 4y ago"Our biggest jobs run MPI, and all pods within the job are participating in a single MPI communicator." https://openai.com/research/scaling-kubernetes-to-7500-nodes https://openai.com/research/scaling-kubernetes-to-7500-nodes
- cavisne 4y agoLLM training (at least the ones we have research details on) sync weights every step so they have very high networking and latency needs. That’s why every cloud vendor competes on interconnect speeds for GPU machines. So they are basically classic supercomputing workloads.
- cavisne 4y agoI don't see a lot of traditional "HPC engineers" in the ml infrastructure space though. I do wonder if in hindsight using a Slurm cluster for scheduling, Lustre for data, MPI for any connectivity thats not covered by NCCL, would have been better than trying to make object storage, grpc, kubernetes, ray etc work.
- the_svd_doctor 4y ago“HPC engineering” and “Optimizing LLM training” are similar. It’s about profiling, performance modeling, finding bottlenecks, rewriting what’s needed, etc. Lots of overlap, and people from more traditional HPC doing it too. Obviously if you’re a 50yo university professor doing airplane simulations you won’t switch to LLMs today though…
- adw 4y agoSome mixture of MPI, NCCL, Gloo, and whatever proprietary stuff TPU clusters do. All of these are basically either trad-HPC in style or literally from the supercomputing community. Interconnects tend to be Infiniband or the like, which, again, straight out of big iron.
- osigurdson 4y agoThey use Kubernetes and MPI https://openai.com/research/scaling-kubernetes-to-7500-nodes https://openai.com/research/scaling-kubernetes-to-7500-nodes
- p_l 4y agoCUDA has support for MPI, including MPI done from GPU itself (over nvlink or with support for accessing host infiniband adapter, iirc).