7 ms·
I'm not sure it was "pointers" but it was some very low cost way to change ownership of memory between the CPU and GPU. They had some fancy marketing name for
by spitfire 2y ago
I'm not sure it was "pointers" but it was some very low cost way to change ownership of memory between the CPU and GPU.
They had some fancy marketing name for it at the time. But it wasn't on all chips, it should have been. Even if it was dog slow between PCIe GPU and CPU the unified interface would have been the right way to go. Also, amenable to automated scheduling.
The point still stand though, I want entirely unified GPU and CPU memory.
- JonChesterfield 2y agoThe unified address space with moving pages between CPU and GPU on page fault works on some discrete GPU systems but it's a bit of a performance risk compared to keeping the pages on the same device. Fundamentally if you've got separate blocks of memory tied together by pcie then it's either annoying copying data across or a potential performance problem doing it behind the scenes. A single block of memory that everything has direct access to is much better. It works very neatly on the APU systems.
- spitfire 2y ago> Fundamentally if you've got separate blocks of memory tied together by pcie then it's either annoying copying data across or a potential performance problem doing it behind the scenes. Well, as I said that's amenable to automated planning. But what I really, really want is a nice APU with 512GB+ of memory that both the CPU and GPU can access willy nilly.
- JonChesterfield 2y agoYep, that's what I want too. The future is now. The MI300A is an APU with 128gb on the package. They come in four socket systems, that's 512gb of cache coherent machine with 96 fast x64 cores and many GCN cores. Quite like a node from El Capitan. I'm delighted with the hardware and not very impressed with the GPU offloading languages for programming it. The GCN and x64 cores are very much equal peers on the machine, the asymmetry baked into the languages grates on me. (on non-apu systems, moving the data around in the background such that the latency is hidden is a nice idea and horrendously difficult to do for arbitrary workloads)
- kcb 2y agoProbably thinking of this https://en.m.wikipedia.org/wiki/Heterogeneous_System_Architecture https://en.m.wikipedia.org/wiki/Heterogeneous_System_Archite... > Even if it was dog slow between PCIe GPU and CPU the unified interface would have been the right way to go That is actually what happened. You can directly access pinned cpu memory over pcie on discrete gpus.