3 ms·
> Another kind of misconception: data transfer is a _really_ overlooked issue. […] If you want to write 20mb of data to a buffer, that's not just a memcpy, all
by ElFitz 2y ago
> Another kind of misconception: data transfer is a _really_ overlooked issue. […] If you want to write 20mb of data to a buffer, that's not just a memcpy, all that data has to go over the PCIe buss to the GPU […], and that's going to be expensive (in real time contexts). Similarly if you want to read a whole large buffer of results back from the GPU, that's going to take some time.
Does having a unified memory, like Apple’s M-series chips, help with that?
- b3orn 2y agoIn theory yes, because you wouldn't need to copy the data, in practice it depends on the API and you might end up copying data from RAM to RAM. If the API doesn't allow you to simply pass an address to the GPU then you need to allocate memory on the GPU and copy your data to that memory, even if it's unified memory.
- Kon-Peki 2y agoFor Apple specifically, you have to act as if you do not have unified memory because Apple still supports discrete GPUs in Metal and also Swift is reference counted - the CPU portion of the app has no idea if the GPU portion is still using something (remember that the CPU and GPU are logically different devices even when they are on the same die). When you are running your code on an M- or A-series processor, most of that stuff probably ends up as no-ops. But worse case is that you copy from RAM to RAM, which is extraordinarily faster than pushing anything across the PCIe bus.
- ElFitz 2y agoGood to know, thanks!