3 ms·
When training over multiple GPUs, it's hard not to think about Ray (https://docs.ray.io/en/latest/train/train.html https://docs.ray.io/en/latest/train/train.htm
by vvipgupta 4y ago
When training over multiple GPUs, it's hard not to think about Ray (https://docs.ray.io/en/latest/train/train.html https://docs.ray.io/en/latest/train/train.html). Ray, as an open-source project, has exploded over the last few years and helps with the memory bottleneck by segregating memory and computing.
FYI, I am not affiliated with Ray. However, I did write the following paper on scaling data-parallel training for large ML models ;)
https://openreview.net/pdf?id=rygFWAEFwS https://openreview.net/pdf?id=rygFWAEFwS
Also, another one of my papers talks about distributed training while reducing the communication bottleneck for distributed training:
https://dl.acm.org/doi/pdf/10.1145/3447548.3467080 https://dl.acm.org/doi/pdf/10.1145/3447548.3467080
- robertnishihara 4y agoI'm one of the Ray developers, thanks for the shoutout :) If you're curious about how Ray is used for LLMs, here are some interesting examples of LLM projects using Ray! - Alpa does training and serving with 175B parameter models https://github.com/alpa-projects/alpa https://github.com/alpa-projects/alpa - GPT-J https://github.com/kingoflolz/mesh-transformer-jax https://github.com/kingoflolz/mesh-transformer-jax - Another HN thread on training LLMs with Ray (on TPUs in this case) https://news.ycombinator.com/item?id=27731168 https://news.ycombinator.com/item?id=27731168 - OpenAI fireside chat on the evolution of their infrastructure and usage of Ray for training https://www.youtube.com/watch?v=CqiL5QQnN64 https://www.youtube.com/watch?v=CqiL5QQnN64 - Cohere on their architecture for training LLMs https://www.youtube.com/watch?v=For8yLkZP5w&t=3s https://www.youtube.com/watch?v=For8yLkZP5w&t=3s Some other thoughts 1. There is a lot more we want to do to make Ray better for working with large language models and for making training, serving, and batch inference work well out of the box. 2. The original post is about training, but we actually see even more interest in fine-tuning and serving with LLMs, in part because there are good pre-trained models. 3. For LLMs, we see a lot of interest in Ray + Jax or Ray + TPUs relative to what we see in other use cases.
- eternalban 4y agoDo you see any convergence on wire (arrow?) and storage (pandas?) formats?
- pavelstoev 4y agoAnd we can make Ray more efficient by optimizing GPU hardware utilization https://centml.ai/ https://centml.ai/
- q1w2 4y agoWill it work with a PC that has 7 AMD Vega GPUs?
- robertnishihara 4y agoYes, but this will largely come down to whether the deep learning framework that you're using (PyTorch, TensorFlow, Jax, etc) works well in that setting. Ray is pretty framework and hardware agnostic and can be used to schedule / scale different ML frameworks on different types of devices (CPUs, GPUs, TPUs, etc), but the actual logic for running code on the accelerators lives in the deep learning framework.