4 ms·
The distinction is that larger installations cannot form a single network. Before xAI's new network architecture, only around 30k GPUs could train a model simul
by enslavedrobot 2y ago
The distinction is that larger installations cannot form a single network. Before xAI's new network architecture, only around 30k GPUs could train a model simultaneously. It's not clear how many can train together with xAI's new approach, but apparently it is >100k.