5 ms·
what kind of bandwidth/latency between GPUs would one need in that setup to not be bottlenecking? What you're describing sounds quite forgiving. Is it forgivi
by dogcomplex 2y ago
what kind of bandwidth/latency between GPUs would one need in that setup to not be bottlenecking? What you're describing sounds quite forgiving. Is it forgiving enough that we could potentially connect those GPUs over a LAN, or even a remote decentralized cloud of host computers?
From my understanding that's certainly possible to do without the latency hurting much with large batching between inference layers
- ryao 2y agoThis depends on: * the model dimension * how many bits per variable are used by your quantization * how many tokens are being processed per step (input processing can do all N input tokens simultaneously and output processing can only do 1 at a time when doing a single query) * how many times you split the model layers across multiple GPUs The model dimensions are: * 4096 for llama 3/3.1 8B. * 8192 for llama 3/3.1 70B. * 16384 for llama 3.1 405B. The model layers are: * 32 for llama 3/3.1 8B * 80 for llama 3/3.1 70B * 126 for llama 3.1 405B The amount of data that needs to be transferred for each split is surprisingly small. Each time you move the calculation of a subsequent layer to a different GPU, you need to transfer an array that is of size model_dimension * num_tokens * bits_per_variable. Then this reduces to a classic network transfer time problem, where you consider both time until the first byte arrives and the transfer time until the last byte arrives. Reality will likely be longer than that idealized scenario, especially since you need to send a signal saying to begin computing. Input processing can tackle so many tokens simultaneously that it probably is not worth thinking too much about this penalty there. Output processing is where the penalty is more significant, since you will incur these costs for every token. Let’s say we are doing fp16 or bf16 on llama 3 8B. Then we need to transfer 8KB every time we move the calculation for another layer to another GPU. If you use RDMA and do this over 10GbE, the transfer time would be 6.4 microseconds. If we assume the time to first byte and time to do a signal to begin processing is 3.6 microseconds combined (chosen to round things up), then we get a penalty of 10 microseconds per split, per token. If you are doing 60 tokens per second and split things across 4 GPUs over the network, you have a penalty of 30 microseconds per token. It will run about 0.003% slower and you are not going to notice this at all. Assuming 10GbE with RDMA is somewhat idealized, although I needed to pick something to give some numbers. In any case, the equation for figuring what factor slower it would be is 1 / (1 + time to do transfers and trigger processing per each token in seconds). That would mean under a less ideal situation where the penalty is 5 milliseconds per token, the calculation will be ~0.99502487562 times what it would have been had it been done in a hypothetical single GPU that has all of the VRAM needed, but otherwise the same specifications. This penalty is also not very noticeable. In conclusion, you are right. I actually recall seeing a random YouTube video talking about a software project that does clustered interferencing, so people are already doing this. Unfortunately, I do not remember the project name, channel name or video name.