4 ms·
The ability to get huge amounts of VRAM/$ is what I find incredibly interesting. A lot of diffusion techniques are incredibly VRAM intensive and high VRAM consu
by chickenpotpie 3y ago
The ability to get huge amounts of VRAM/$ is what I find incredibly interesting. A lot of diffusion techniques are incredibly VRAM intensive and high VRAM consumer cards are rare and expensive. I'll gladly take the slower speeds of an APU if it means I can load the entire model in memory instead of having to offload chunks of it.
- dragontamer 3y agoUsing DDR5 as VRAM however means you're only getting 50GB/s to 100GB/s read/write speed instead of the 500GB/s available on a proper GPU. That might be a fine tradeoff for some kernels. But my understanding is that Stable Diffusion is very VRAM-bandwidth heavy and actually benefits from the higher-speed GDDR6 or HBM RAM on a proper high-end GPU.
- thulle 3y agoWith a 4090 nvidia-smi reports ~60-70% mem bandwidth usage while at 99% gpu usage, so that'd be 650-750GB/s. Considering how slow the inference is on the APU having 10% of the 4090 mem bandwidth maybe isn't that much of an issue? edit: comment from the reddit thread: > LLM isn't all that great as it is primarily memory bandwidth bound, ie almost no difference from a CPU if your memory bw is mere 12/25Gb/s. SD needs far more compute for inference - APU with slow memory helps. - https://old.reddit.com/r/Amd/comments/15t0lsm/i_turned_a_95_amd_apu_into_a_16gb_vram_gpu_and_it/jwjr5j3/ https://old.reddit.com/r/Amd/comments/15t0lsm/i_turned_a_95_...
- delusional 3y agoA PCIE 4.0 16x link should provide around 32GB/s of bandwidth, close to the 33GB/s of DDR5-3200. In a perfect world, it would seem to me that doing 100% offloading (streaming everything as needed from system memory) should be equivalent to doing the calculations in system memory in the first place. The GPU memory just acts as a cache, and should only speed up the processing.
- BobbyJo 3y agoThe full mem-swap congo line is DRAM<->PCIE<->VRAM<->GPU. PCIE is the weak link, but I'd be willing to bet that those transfers are't 100% overlapped, and PCIE transfer rate represents best case speed as opposed to expected. In the case of unidirectional writes, you'd have to cut that speed in half.
- dragontamer 3y agoEvery modern motherboard is dual-channel DDR5 at a minimum, maybe quad-channel. 2x sticks of 32GB/s RAM, properly configured, will run at 64GB/s of bandwidth. Modern servers are quad, hex, or oct-channel (4x, 6x, or 8x parallel sticks of RAM) in practice. Or even more (ex: Intel Xeon Platinums are 6x channel per socket, so an 8x socket 8x CPU Xeon Platinum will be like 48x DDR5 parallel RAM sticks). ---------- PCIe x16 by the way, is 16x parallel lanes of PCIe. All the parallelism is already innate in modern systems. --------- L3 cache is TB/s bandwidth IIRC. CPUs inside of CPU-space will automatically be caching a lot of those RAM commands, so you can go above the RAM-bandwidth limitations in practice, though it depends on how your code accesses RAM. GPUs have very small caches, and have higher latency to those caches. Instead, GPUs rely upon register-space and extreme amounts of SMT/wavefronts to kinda-sorta hyperthread their cores to hide all that latency.
- gautamcgoel 3y agoThis argument makes sense only if you assume that your model is too large to fit in the GPU's RAM and hence has to reside mainly in the CPU's RAM.
- delusional 3y agoThat's true. The comment i was responding to talked about offloading. I was assuming he was talking about offloading part of the core model to the system RAM which would need to be reloaded frequently.
- justinclift 3y agoIf you're ok with older generation Nvidia gear, but still with 24GB ram, then some people are using things like this: https://www.ebay.com.au/itm/126046615075 https://www.ebay.com.au/itm/126046615075 That's $250 in Australian dollars though, which is about US$160. I'm not affiliated with that seller btw, I just remembered the search result from looking a while back. :)
- cptskippy 3y agoI bought one of those earlier this week and I'm in the process of getting it setup.
- choppaface 3y agoThese have lots of memory but are pretty slow, can be 50x-100x slower than a gamer card from the past couple years, plus lots of heat / power inefficiency. If you have any software than can take advantage of tensor cores / matrix cores or int8 / fp16 ops then modern hardware will probably win.
- deleted 3y ago[deleted]
- sp332 3y ago(It's a Tesla M40 with a blower fan retrofit, "Buy it Now" price $300 Australian dollars.)
- lhl 3y agoFor not much more (US$200) you can find lots of P40s, which is a generation newer and will give you double the memory bandwidth and FP32. That being said, used 3090s are going for about $600 now and are much better bang/buck and easier (software and hardware) to setup.
- smcleod 3y agoHey fellow person-in-Australia, I bought a P100 for $250 on eBay and have been using that with some custom cooling I printed and it works pretty damn well. You wouldn’t want an M or K series Tesla though as they’re just too old and not powerful enough to be that useful. Here’s some photos https://aus.social/@s_mcleod/110841559904867676 https://aus.social/@s_mcleod/110841559904867676
- Auracle 3y agoI mean, it’s ridiculously slower. They’re getting something like .5it/s at 512x512. I believe my 3080 gets 10? Maybe more If AMD improves their speed with this configuration on later APUs though - that could really hurt Nvidia.
- kramerger 3y agoThis is still a 5-10x improvement over the CPU. And most people playing with AI don't have an expensive GPU with 12+ GB VRAM.
- andrewstuart 3y agoYou mean in the Reddit post? That’s using a 4600G.