4 ms·
Rail-Only: A Low-Cost High-Performance Network for Training LLMs with T Params
- teleforce 2y agoPlease check this HN post on the similar subject by Meta [1]. Previous paper by the same team from Meta and MIT but with Billions instead of Trillions of parameters [2]. [1] A RoCE network for distributed AI training at scale: https://news.ycombinator.com/item?id=41162664 https://news.ycombinator.com/item?id=41162664 [2] Optimized Network Architectures for Training Large Language Models With Billions of Parameters [PDF]: https://people.csail.mit.edu/ghobadi/papers/rail_llm_hotnets_2023.pdf https://people.csail.mit.edu/ghobadi/papers/rail_llm_hotnets...