6 ms·
Ask HN: Resources for general purpose GPU development on Apple's M* chips?
While Apple M* chips seems to have an incredible unified memory access, the available learning resources seem to be quite restricted and often convoluted. Has anyone been able to get past this barrier?
I have some familiarity with general purpose software development with CUDA and C++. I want to figure how to work with/ use Apple's developer resources for general purpose programming.
- barkingcat 2y agoThere is no general purpose GPU development on Apple M series. There is Metal development. You want to learn Apple M-series gpu and gpgpu development? Learn Metal! https://developer.apple.com/metal/ https://developer.apple.com/metal/
- kristianp 2y ago> There is no general purpose GPU That's what GPGPU stands for. So your 2 sentences contradict each other.
- rowanG077 2y agoIf you are open to run Linux you can use standard opencl and vulkan.
- deleted 2y ago[deleted]
- feznyng 2y agoBesides the official docs you can check out llama.cpp as an example that uses metal for accelerated inference on Apple silicon.
- dylanowen 2y agoPeople have already mentioned Metal, but if you want cross platform, https://github.com/gfx-rs/wgpu https://github.com/gfx-rs/wgpu has a vulkan-like API and cross compiles to all the various GPU frameworks. I believe it uses https://github.com/KhronosGroup/MoltenVK https://github.com/KhronosGroup/MoltenVK to run on Macs. You can also see the metal shader transpilation results for debugging.
- grovesNL 2y agowgpu has its own Metal backend that most people use by default (not MoltenVK). There is also a Vulkan backend if you want to run Vulkan through MoltenVK though.
- rudedogg 2y agoWith what the OP asked for, I don't think wgpu is the right choice. They want to push the limits of Apple Silicon, or do Apple platform specific work, so an abstraction layer like wgpu is going in the opposite direction in my opinion. Metal, and Apple's docs are the place to start.
- PittleyDunkin 2y agoIndeed. I'm curious how much overhead there is in practice given the fact that the hardware wasn't designed to provide vulkan support. I honestly have no clue what to expect.
- morphle 2y agoYou can help with the reverse engineering of Apple Silicon done by a dozen people worldwide, that is how we find out the GPU and NPU instructions[1-4]. There is over 43 trillion float operations per second to unlock at 8 terabit per second 'unified' memory bandwidth and 270 gigabits per second networking (less on the smaller chips).... [1] https://github.com/AsahiLinux/gpu https://github.com/AsahiLinux/gpu [2] https://github.com/dougallj/applegpu https://github.com/dougallj/applegpu [3] https://github.com/antgroup-skyward/ANETools/tree/main/ANEDisassembler https://github.com/antgroup-skyward/ANETools/tree/main/ANEDi... [4] https://github.com/hollance/neural-engine https://github.com/hollance/neural-engine You can use a high level APIs like MLX, Metal or CoreML to compute other things on the GPU and NPU. Shadama [5] is an example programming language that translates (with Ometa) matrix calculations into WebGPU or WebGL APIs (I forget which). You can do exactly the same with the MLX, Metal or CoreML APIs and only pay around 3% overhead going through the translation stages. [5] https://github.com/yoshikiohshima/Shadama https://github.com/yoshikiohshima/Shadama I estimate it will cost around $22K at my hourly rate to completely reverse engineer the latest A16 and M4 CPU (ARMV9), GPU and NPU instruction sets. I think I am halfway on the reverse engineering, the debugging part is the hardest problem. You would however not be able to sell software with it on the APP Store as Apple forbids undocumented API's or bare metal instructions.
- KeplerBoy 2y agoWhere does the 270 gbit/s networking figure come from? Is it the aggregate bandwidth from the pcie slots on the mac pro, which could support nics at that speeds (and above according to my quick maths#), but there is not really any driver support for modern Intel or Mellanox/Nvidia NICs as far as I can tell. My use case would be hooking up a device which spews out sensor data at 100 gbit/s over qsfp28 ethernet as directly to a GPU as possible. The new mac mini has the GPU power, but there's no way to get the data into it. # 2x Gen4x16 + 4x Gen3x8 = 2 * 31.508 GB/s + 4 * 7.877 GB/s ≈ 90 GB/s = 720 gbit/s
- morphle 2y ago> Where does the 270 gbit/s networking figure come from? Is it the aggregate bandwidth from the pcie slots on the Mac pro We both should restate and specify the calculation for each different Apple Silicon chip and the PCB/machine model it is wired onto. The $599 M4 Mac mini base model networking (aggregated Wifi, USB-C, 10G Ethernet, Thunderbolt PCIe) is almost 270 Gbps. Your 720 Gbps is for a >$8000 Mac Pro M2 Ultra but the number is to high because the 2x Gen4x16 is shared/oversubscribed with the other PCIe lanes for x8 PCIe slots, SSD and Thunderbolt. You need to measure/benchmark it, not read the marketing PR. I estimate the $1400 M4 Pro Mac mini networking bandwidth by adding the external WiFi, 10 Gbps Ethernet, two USC-C ports (2 x 10 Gbps) and three Thunderbolt 4 ports (3 x 80/120 Gbps) but subtracting the PCIe 64 Gbps limit and not counting the internal SSD. Two $599 M4 Mac mini base models are faster and cheaper than one M4 Pro Mac mini. The point of the precise actual measurements I did of the trillion opereations per second and the billion of bits per second networking/interconnect of the M4 Mac mini against all the other Apple silicon machines is to find which package (chip plus pcb plus case) has the best price/performance/watt balanced against them networked together. On januari 2025 you can build the cheapest fastest supercomputer in the world from just off the shelf M4 16Gb Mac mini base models with 10G Ethernet, Mikrotek 100G switches and a few FPGA's. It would outperform all Nvidia, Cerebras, Tenstorrent and datacenter clusters I know of, mainly because of the low power Apple Silicon. Note that the M4 has only 1,2 Tips unified memory bandwidth and the M4 Pro has double that. The 8 Tops unified memory bandwidth is on the M1 and M2 Studio Ultra with 64/128/192GB DRAM. Without it you cant's reach 50 trillion operations per second. A Mac Studio has only around 190 Gbps external networking bandwidth but does not reach 43 trillion TOPS, as does the 720 Gbps of your Mac Pro estimate. By reverse engineering the instruction set you could squeeze a few percent extra performance out of this M4 cluster. The 43 trillion TOPS of the M4 itself is an estimate. The ANE does 34 TOPS, the CPU less than 5 TOP depending on float type and we have no reliable benchmarks for the CPU floating point.
- mkagenius 2y agoCheck out MLX[1]. Its a bit like pytorch/tensorflow with added benefit of Apple Silicon. 1. https://ml-explore.github.io/mlx/build/html/index.html https://ml-explore.github.io/mlx/build/html/index.html
- rgovostes 2y agoIt's hard to answer not knowing exactly what your aim is, or your experience level with CUDA and how easily the concepts you know will map to Metal, and what you find "restricted and convoluted" about the documentation. <Insert your favorite LLM> helped me write some simple Metal-accelerated code by scaffolding the compute pipeline, which took most of the nuisance out of learning the API and let me focus on writing the kernel code. Here's the code if it's helpful at all. https://github.com/rgov/thps-crack https://github.com/rgov/thps-crack
- nixpulvis 2y ago2024 and still finding cheat codes in Tony Hawk Pro Skater 2. Wild!
- selimthegrim 2y agoIf Jamie Kennedy is reading this, we still haven’t found the cheat code to make you funny.
- thetwentyone 2y agoI’ve had a good time dabbling with Metal.jl: https://github.com/JuliaGPU/Metal.jl https://github.com/JuliaGPU/Metal.jl
- Archit3ch 2y agoSame. It can even run realtime workloads (audio).
- amelius 2y agoApple is known to actively discourage general purpose computing. Better try a different vendor.
- codr7 2y agoPreferably one that sells computers, not fashion statements.
- likeabbas 2y agoIt's not a fashion statement, it's a fucking deathwish
- saagarjha 2y agoidk about “known” considering they basically created OpenGL
- mixmastamyk 2y agoThat was SGI.
- talldayo 2y agoIf OpenGL is your most up-to-date reference for Apple supporting general purpose computing then I think it absolutely emphasizes how little work they've put in.
- aleinin 2y agoIf you're looking for a high level introduction to GPU development on Apple silicon I would recommend learning Metal. It's Apple's GPU acceleration language similar to CUDA for Nvidia hardware. I ported a set of puzzles for CUDA called GPU-Puzzles (a collection of exercises designed to teach GPU programming fundamentals)[1] to Metal [2]. I think it's a very accessible introduction to Metal and writing GPU kernels. [1] https://github.com/srush/GPU-Puzzles https://github.com/srush/GPU-Puzzles [2] https://github.com/abeleinin/Metal-Puzzles https://github.com/abeleinin/Metal-Puzzles
- dylan604 2y agoAfter a quick scan through the [2] link, I have added this to the list of things to look into in 2025
- Jiahang 2y agoCurious about the others in your list
- singlepaynews 2y agoCan anyone recommend a CUDA equivalent of (2)? That’s a spectacular learning resource and I’d like to use a similar one to upskill for CUDA
- dagmx 2y agoIsn’t the link right before it exactly what you’re asking for? Since 2 is a port of 1
- desideratum 2y agoI'd reccomend checking out the CUDA mode Discord server! They also have a channel for Metal https://discord.gg/ZqckTYcv https://discord.gg/ZqckTYcv
- TriangleEdge 2y agoWhy not OpenCL or OpenGL? You'll not be constrained by the flavor of GPU.
- nox101 2y agoSounds like you've never actually tried running those two APis across platforms? if you want portable use WebGPU either via wgpu for rust or dawn for C++ They actually do run on Windows, Linux, Mac, iOS, and Android portably
- thrtythreeforty 2y agowgpu Just Works from C++ as well. Both projects implement the webgpu.h API
- billti 2y agoIf you know CUDA, then I assume you know a bit already about GPUs and the major concepts. There’s just minor differences and different terminology for things like “warps” etc. With that base, I’ve found their docs decent enough, especially coupled with the Metal Shader Language pdf they provide (https://developer.apple.com/metal/Metal-Shading-Language-Specification.pdf https://developer.apple.com/metal/Metal-Shading-Language-Spe...), and quite a few code samples you can download from the docs site (e.g. https://developer.apple.com/documentation/metal/performing_calculations_on_a_gpu https://developer.apple.com/documentation/metal/performing_c...). I’d note a lot of their stuff was still written in Objective-C, which I’m not that familiar with. But most of that is boilerplate and the rest is largely C/C++ based (including the Metal shader language). I just ported some CPU/SIMD number crunching (complex matrices) to Metal, and the speed up has been staggering. What used to take days now takes minutes. It is the hottest my M3 MacBook has ever been though! (See https://x.com/billticehurst/status/1871375773413876089 https://x.com/billticehurst/status/1871375773413876089 :-)