7 ms·
Simplifying GPU Application Development with HMM
- diabllicseagull 3y agothey used to not support unified memory in their vGPU drivers. it was a major deal breaker back then.
- aseipp 3y agoI think that's still the case; though I assume you're talking about Linux-on-Linux vGPUs, the same is true of e.g. WSL2 where unified memory isn't supported. Sucks, because it's a great feature.
- av3csr 3y agoAccording to the manual, UVM is supposed to be working on vGPUs (at least MIG-backed vCS), I could never get it working though
- pjmlp 3y ago> This new ability to directly read or write to the full application memory address space will significantly improve programmer productivity for all programming models built on top of CUDA: CUDA C++, Fortran, standard parallelism in Python, ISO C++, ISO Fortran, OpenACC, OpenMP, and many others. This is the part of CUDA alternatives always miss when their models only support C and some C++ subset.
- mratsim 3y agoApple Metal does this though.
- capableweb 3y agoAlthough Apple holds a <10% of total market share when it comes to computers, so not sure how helpful it is.
- Me1000 3y agoWell Nvidia holds like 90% of the GPU marketshare, so any reply mentioning a competitor would have this property.
- liuliu 3y agoOnly recently. mmap file and directly use in Metal kernels is actually not supported until iOS 16 / macOS 13. Also, there are limited optimization opportunities around that and the recommended way seems still to use the specific Metal APIs to stream load assets from disk.
- pjmlp 3y agoYou missed the polyglot description regarding which workloads CUDA supports.
- viraptor 3y agoDoes that mean that now anyone with rtx20 series or above can run local ML models as big as their RAM allows? (Or larger if they're happy to wait for swapping to SSD) Or am I misunderstanding the scale of the impact here? (Not exactly "now", but when the software is recompiled / ported to this)
- smcleod 3y agoYou can already do that with GGUF/GGML models which allow you to split between CPU and GPU. Obviously there is a performance hit when running on your DDR5 and CPU compared to HBM/GDDR and GPU but it’s better than nothing.
- 0cf8612b2e1e 3y agoI have not been keeping up with developments. Does this mean mortals can run the biggest tier of Llama models (albeit with trash performance) by using system ram? For playing around, I would be willing to let my system chug along just to see what the top tier models can achieve.
- smcleod 3y agoTechnically yes - if you have lots of ram you can use that and your CPU, as you say, the performance would be pretty poor, though, especially as it’s a toll where you want to tweak your responses quite frequently. I’ve been running and old Nvidia Tesla P100 card. I got cheap on eBay for awhile now it has 16 GB of VRAM but it is pretty old. I’m so interested in this now I’ve gone out and got myself a secondhand RTX 3090 - something I never thought I’d do, but I’d really like to run 30B models in GPU.
- _w1tm 3y agoYes. I recently benchmarked the 70B Llama 2 model on a 24 vCPU vSphere host with 64GB RAM (through Ollama) and it was capable of spitting out ~0.15 tokens / second. Useless for any interactive use-case but better than nothing. As a comparison the 7B Llama 2 model was ~1.5 tokens / second on the same hardware while the cheapest M1 MacBook Air can do ~10 tokens / second thanks to GPU acceleration.
- peter_d_sherman 3y ago>"As an aside, new hardware platforms such as NVIDIA Grace Hopper natively support the Unified Memory programming model through hardware-based memory coherence among all CPUs and GPUs. For such systems, HMM is not required, and in fact, HMM is automatically disabled there. One way to think about this is to observe that HMM is effectively a software-based way of providing the same programming model as an NVIDIA Grace Hopper Superchip." 1) I am curious what the AMD equivalent of nVidia's HMM is, or will be... 2) I am curious if software will be able to be written with HMM (or some higher level abstraction API) such that HMM enabled software will also function on an AMD or other 3rd party GPU...
- dagmx 3y agoAMDs answer will be “nothing” imho. They’ve really left this area wide open for over a decade now when it’s been extremely clear this is where the market was going. Their GPU and GPU compute story is a mess, because rocm has the most confusing compatibility story possible . They’ve been late to compute accelerators as well. I don’t think there’ll be any abstraction layers either. The community as a whole is more than happy to be single vendor. AMD has shown they can’t build compute stacks, not because of technology reasons but purely long term decisions. The community therefore won’t do it for them.
- kimixa 3y agoROCm already supports HMM. You're not helping anything by going off on some rant based on an assumption and falsehood - this sort of comment is exactly the sort of thing the phrase "FUD" is used to describe.
- dagmx 3y agoYou’re right that my rant is incorrect on the premise that they don’t have hmm, but it’s because I missed rocm adding it two years ago. So my bad, and unfortunately I can’t edit my post so I’ll leave the link here with my apologies. https://www.phoronix.com/news/Radeon-ROCm-4.3 https://www.phoronix.com/news/Radeon-ROCm-4.3 The reason I missed it is because rocm dropped support for my cards very unceremoniously. At which point I gave up. I do think the rest of my point outside of the first sentence is valid though. Rocm isn’t reliable to target. Nowhere near CUDA. That it’s so dependent on what card you have, what OS/kernel you use and is so aggressive with dropping support for older cards, makes the entire ecosystem a mess. CUDA by comparison is so much more ubiquitous. That becomes chicken and egg with popular libraries adding rocm support because it then ends up targeting such a sliver (and shifting sliver) at that of the market.
- amelius 3y ago"What every programmer should know about memory" needs an update. https://people.freebsd.org/~lstewart/articles/cpumemory.pdf https://people.freebsd.org/~lstewart/articles/cpumemory.pdf
- dragontamer 3y agoI don't think so. The only thing that's been added is bank-groups in DDR4 IMO. But all you need to know is that modern RAM is maybe 16x to 32x way parallel per stick. The interface operates are faster than RAM can respond in time, so an "Optimal" CPU will list off 32x to 64x (32x for the first stick, 32x for the 2nd stick) read/write commands before the first command ever responds. Understanding that mechanism is what that document is about (how CPUs coalesce memory and parallelizes requests). ---------------- GPUs have one additional coalesce layer given channel vs bank conflicts, and all that noise. But most GPU manuals (be they NVidia or AMD) will cover those details.
- drewg123 3y agoI used to work on Myrinet HPC NICs many years ago, and the ability for a PCI(e) device to access any memory by user virtual address was a desirable feature. I believe that Quadrics did this first using a patched version of DEC OSF/1 (UNIX, Tru64, whatever you want to call it), where they hooked into the kernel pmap (page table) code, and sync'ed the page tables with their NIC. That way the NIC could do the virtual to physical translations, and know if a virtual memory address was backed by a physical page. What Nvidia is doing here sounds similar. Does linux provide such primitives now?
- drewg123 3y agoIts really hard to google for information on older stuff like this. I did find a presentation from 2000 where they talk about "OS Bypass with Virtual Addressing; no page locking or copying; full protection" (https://hsi.web.cern.ch/HNF-Europe/sem3_2001/hnf.pdf https://hsi.web.cern.ch/HNF-Europe/sem3_2001/hnf.pdf)
- jhj 3y agoFor performance, it's always better to explicitly manage GPU memory and host/device copies for performance than to depend upon the unified memory paging mechanism, if it's possible to go the extra effort. My feeling is that unified memory and on-demand paging introduced with Pascal? was mainly about making it easier to onboard existing applications (e.g., HPC codes etc) to the GPU a bit at a time with less problem. For writing a GPU application from scratch, I don't think it makes much sense (unless the granluarity of the data that you are moving around is really tiny and/or you can't predict what you would need in advance on CPU or GPU).
- deleted 3y ago[deleted]