4 ms·
It can make a difference when using tensor parallelism to run small batch sizes. Not a huge difference like training because we don't need to update all weights
by andersa 2y ago
It can make a difference when using tensor parallelism to run small batch sizes. Not a huge difference like training because we don't need to update all weights, but still a noticeable one. In the current inference engines there are some allreduce steps that are implemented using nccl.
Also, paged KV cache is usually spread across GPUs.