4 ms·Efficiently Scale LLM Training Across a Large GPU Cluster with Alpa and Ray1 points by dmatrixjsd 3y ago