3 ms·
Yeah, I'm sure the performance of this must be pretty bad. As I understand it, non-batched LLM inference is generally limited by the bandwidth required to read
by makomk 3y ago
Yeah, I'm sure the performance of this must be pretty bad. As I understand it, non-batched LLM inference is generally limited by the bandwidth required to read the entire weights from GPU memory for each token produced. This trick instead moves the bandwidth bottleneck to loading those weights into the GPU from wherever has enough space to store them, probably some kind of SSD, and that has much lower bandwidth than GPU RAM.
- andy99 3y agoFor CPU inference, performance is generally memory bandwidth limited. A modern cpu can do the matmuls (all of which are matrix-vector) faster than the data can be loaded in from memory. If the model is all in RAM and being moved one layer at a time to a GPU, would that be faster than getting it to the CPU? What about if the model is on an SSD. The point being, compute is not the bottleneck so I wonder if this provides speedup compared to an optimized model run on a CPU?
- faeriechangling 3y agoIt would potentially be slower because pcie bandwidth is likely to be lower than memory bandwidth. The way to get a speed up is to split the load across both the CPU and GPU since LLMs can be split into layers. This also increases the size the model can be.
- andy99 3y agoJust want to add, this approach is the exact opposite of conventional wisdom around cache hierarchy. Here we are loading data in from the slowest source, using it once, and throwing it away. I'm not sure what can be changed with a sequential model like and llm, just that this is almost the textbook slowest way to run the model.
- gpm 3y agoTried this awhile ago on the llama models, with my hardware 1. The GPU was still faster, though the difference was smaller than I really expected. 2. Almost all the time was spent on compute, not loading weights to/from the GPU. For models I could fit on my GPU the difference in performance in loading one layer at a time vs loading them all was completely negligible, and the difference in memory usage (constraining what else I could do with the computer simultaneously) significant. 3. Loading from SSD was a huge bottleneck when I tried that Hardware: RTX 2070 super, 5800x
- lmeyerov 3y agoIt sounds like they do it naively, which is fair for a POC Even for non-batched, you can still optimize this a bunch for fairly big models without much more work. Basically, overlap io & compute: while one layer is computing, already be loading in the next layer(s) in parallel. PCI speeds are in the 1GB/s-32GB/s each way for consumer cards nowadays. Without a lot of work, they might be able to fully hide the IO latency.. very model + hw dependent.