6 ms·
It depends how much, where, and when you'd like the data to be sent? Most discrete GPUs cannot access the CPU memory directly, so you need to make a copy throug
by fathyb 4y ago
It depends how much, where, and when you'd like the data to be sent? Most discrete GPUs cannot access the CPU memory directly, so you need to make a copy through the PCI bus, which can be slow.
If you have small chunks of data (mining), or your destination is the screen (video game), it might make sense to use the GPU. If you need high-throughput, low latency, or your destination is something like a sound card (DAW), then SIMD might be a better choice.
Fast integrated GPUs like Apple's allow for directly accessing the main memory without copy, making the GPU more viable for general purposes.
- the__alchemist 4y agoThat makes sense; thanks! And for passing to GPU, you need to serialize the data as byte arrays with specific alignment requirements that are not intuitive.
- Kon-Peki 4y ago> Fast integrated GPUs like Apple's allow for directly accessing the main memory without copy, making the GPU more viable for general purposes. My understanding (and I would be very happy for any clarifications/corrections!) is that you must use Metal Buffers (MTLBuffer) with the Apple Silicon GPU. If your data isn't already in a Metal Buffer (why would it be?), you have to do a copy into a buffer (but it will be extremely fast). Metal Buffers use raw bytes so they will work well with C/C++ arrays, but if you are using something like a Swift array, that is not such an easy to use pairing - you can only get a Swift UnsafeMutableRawPointer to the data. Further complicating things is that for the best GPU compute performance on Apple Silicon, you want to use Metal Buffers that are private to the GPU, forcing you to use a "blit" operation to copy data to/from CPU-visible memory. I've been exploring general purpose computing with both CUDA and Apple Silicon and am having a lot better luck using SIMD intrinsics instead. While the GPU compute is incredibly fast, there is just so much time spent synchronizing data. There really seems to be some sweet spot where the GPU wins out, but in a lot of general purpose uses you won't ever get there. But I would really like to be shown that I'm wrong!!! PS - the very basic ARM Neon stuff that the M1 supports is insanely fast and gets even better when you use multiple threads.
- fathyb 4y ago> My understanding (and I would be very happy for any clarifications/corrections!) is that you must use Metal Buffers (MTLBuffer) with the Apple Silicon GPU. If your data isn't already in a Metal Buffer (why would it be?), you have to do a copy into a buffer (but it will be extremely fast). You have to use Metal Buffers, but you don't have to copy, as long as it is properly page aligned [0]. The YUV data from both camera is for example. > PS - the very basic ARM Neon stuff that the M1 supports is insanely fast and gets even better when you use multiple threads. It is, these chips have so much to give! [0]: https://developer.apple.com/documentation/metal/mtldevice/1433382-newbufferwithbytesnocopy https://developer.apple.com/documentation/metal/mtldevice/14...
- Kon-Peki 4y agoI have just been nerd sniped. EDIT - any hints on how to get a plain-old Swift array to be allocated in a conforming alignment (4096 bytes on Apple Silicon)?
- saagarjha 4y agoIt's not supported, unfortunately. But you can call posix_memalign and work with it as an UnsafeBufferPointer with a similar API.
- shihab 4y ago> I've been exploring general purpose computing with both CUDA and Apple Silicon and am having a lot better luck using SIMD intrinsics instead What sort of programs are you trying this with?
- Kon-Peki 4y agoI am writing some stuff that will hopefully be published next year, so I don't want to get into it too much now. But it basically started with this sequence of text from "Is Parallel Programming Hard, And, If So, What Can You Do About It?" [0]: > Parallel programming has earned a reputation as one of the most difficult areas a hacker can tackle. > However, new technologies that are difficult to use at introduction invariably become easier over time. > Therefore, if you wish to argue that parallel programming will remain as difficult as it is currently perceived by many to be, it is you who bears the burden of proof We are now in an era in which is it nearly impossible for the average consumer to buy a computing product that doesn't have multiple cores, SIMD, and a GPGPU. So I felt that it was time to explore how to do it with everyday basic general computing tasks. I was starting with nearly zero experience and writing what I've been learning :) By the way, even the Raspberry Pi 4 does very nicely with ARM Neon, though it doesn't improve with multiple threads concurrently executing SIMD code like Apple Silicon does! [0] https://mirrors.edge.kernel.org/pub/linux/kernel/people/paulmck/perfbook/perfbook.html https://mirrors.edge.kernel.org/pub/linux/kernel/people/paul...