3 ms·
I remember years ago one of the amd apus had the cup and gpu on the same die, and could exchange ownership of cpu and gpu memory with just a pointer change or s
by spitfire 2y ago
I remember years ago one of the amd apus had the cup and gpu on the same die, and could exchange ownership of cpu and gpu memory with just a pointer change or some other small accounting.
Has this returned? Because for dual gpu/cpu workloads (alpha zero, etc) that would deliver effective “infinite bandwidth” between gpu and cpu. Using an apu of course gets you huge amounts of slowish memory. But being some to fling things around with abandon would be an advantage, particularly for development.
- wmf 2y agoI assume the MI300A APU also supports zero-copy. Because MI300X is a separate chip you necessarily have to copy data over PCIe to get it into the GPU.
- rbanffy 2y agoOne day someone will build a workstation around that chip. One day…
- JonChesterfield 2y agoYou don't need to change the pointer value. The GPU and the CPU have the same page table structures and both use the same pointer representation for "somewhere in common memory". On the GPU there are additional pointer types for different local memory, e.g. LDS is a uint16_t indexing from zero. But even there you can still have a single pointer to "somewhere" and when you store to it with a single flat addressing instruction the hardware sorts out whether it's pointing to somewhere in GPU stack or somewhere on the CPU. This works really well for tables of data. It's a bit of a nuisance for code as the function pointer is aimed at somewhere in memory and whether that's to some x86 or to some gcn depends on where you got the pointer from, and jumping to gcn code from within x86 means exactly what it sounds like.
- spitfire 2y agoI'm not sure it was "pointers" but it was some very low cost way to change ownership of memory between the CPU and GPU. They had some fancy marketing name for it at the time. But it wasn't on all chips, it should have been. Even if it was dog slow between PCIe GPU and CPU the unified interface would have been the right way to go. Also, amenable to automated scheduling. The point still stand though, I want entirely unified GPU and CPU memory.
- JonChesterfield 2y agoThe unified address space with moving pages between CPU and GPU on page fault works on some discrete GPU systems but it's a bit of a performance risk compared to keeping the pages on the same device. Fundamentally if you've got separate blocks of memory tied together by pcie then it's either annoying copying data across or a potential performance problem doing it behind the scenes. A single block of memory that everything has direct access to is much better. It works very neatly on the APU systems.
- spitfire 2y ago> Fundamentally if you've got separate blocks of memory tied together by pcie then it's either annoying copying data across or a potential performance problem doing it behind the scenes. Well, as I said that's amenable to automated planning. But what I really, really want is a nice APU with 512GB+ of memory that both the CPU and GPU can access willy nilly.
- JonChesterfield 2y agoYep, that's what I want too. The future is now. The MI300A is an APU with 128gb on the package. They come in four socket systems, that's 512gb of cache coherent machine with 96 fast x64 cores and many GCN cores. Quite like a node from El Capitan. I'm delighted with the hardware and not very impressed with the GPU offloading languages for programming it. The GCN and x64 cores are very much equal peers on the machine, the asymmetry baked into the languages grates on me. (on non-apu systems, moving the data around in the background such that the latency is hidden is a nice idea and horrendously difficult to do for arbitrary workloads)
- kcb 2y agoProbably thinking of this https://en.m.wikipedia.org/wiki/Heterogeneous_System_Architecture https://en.m.wikipedia.org/wiki/Heterogeneous_System_Archite... > Even if it was dog slow between PCIe GPU and CPU the unified interface would have been the right way to go That is actually what happened. You can directly access pinned cpu memory over pcie on discrete gpus.