3 ms·
The reason is the number of parallel instructions you need to run during training. Let’s say you have a 32 core CPU. Well great. But an A100 GPU has 6912 CUD
by binarymax 4y ago
The reason is the number of parallel instructions you need to run during training. Let’s say you have a 32 core CPU. Well great. But an A100 GPU has 6912 CUDA cores and 432 Tensor cores.
A 32 core CPU may have more capability per core, but it doesn’t scale to meet the needs of training LLMs