3 ms·
I was browsing r/localllm recently and was surprised that some people were purchasing two 3060 12GB and using them in tandem somehow. I actually didn't think th
by reginald78 2y ago
I was browsing r/localllm recently and was surprised that some people were purchasing two 3060 12GB and using them in tandem somehow. I actually didn't think this would work at all without nvlink.
- magicalhippo 2y agoYou just copy data over PCIe, however it is slower than nvlink. There are several ways to utilize multiple GPUs, the main contenders as I understand it is pipelined and so-called tensor parallelism. Think of it as slicing a loaf of bread in regular slices vs along its longer axis. The former can have higher peak throughput while the latter can have lower latency, though it depends on the details[1]. [1]: https://blog.squeezebits.com/vllm-vs-tensorrtllm-9-parallelism-strategies-36310 https://blog.squeezebits.com/vllm-vs-tensorrtllm-9-paralleli...
- ai-christianson 2y ago> You just copy data over PCIe AFAIK this only happens directly over PCIe if using hacked drivers with p2p enabled (I think tinygrad/tinybox provided these drivers initially.) Otherwise, data goes through system bus/CPU first. `nvidia-smi topo -m` will show how the GPUs are connected.
- ai-christianson 2y ago> I actually didn't think this would work at all without nvlink. It will work, but with less memory bandwidth. If using hacked drivers with p2p enabled, it will depend on the PCIe topology. Otherwise, the data will take a longer path. Depending on the model, it may not be that big of a hit in performance. Personally, I have seen higher perf on nvlinked GPUs and models that can fit on 2x 3090 24gbs.
- reginald78 2y agoYes, no one seemed to be bragging about much of a performance boost. I think they were happy being able to run models (at any speed) that needed 24GB without having to buy a 3090.