24 ms·
Apple's MLX adding CUDA support
- gsibble 1y agoAwesome
- nxobject 1y agoIf you're going "wait, no Apple platform has first-party CUDA support!", note that this set of patches also adds support for "Linux [platforms] with CUDA 12 and SM 7.0 (Volta) and up". https://ml-explore.github.io/mlx/build/html/install.html https://ml-explore.github.io/mlx/build/html/install.html
- teaearlgraycold 1y agoI wonder if Jensen is scared. If this opens up the door to other implementations this could be a real threat to Nvidia. CUDA on AMD, CUDA on Intel, etc. Might we see actual competition?
- _ea1k 1y agoI think this is the other way around. It won't be cuda on anything except for nvidia. However, this might make mlx into a much stronger competitor for Pytorch.
- teaearlgraycold 1y agoOh bummer. Almost got excited.
- baby_souffle 1y agoIf you implement compatible apis, are you prohibited from calling it cuda?
- moralestapia 1y agoI'm sure I saw this lawsuit somewhere ... The gist is the API specification in itself is copyright, so it is copyright infringement then.
- deleted 1y ago[deleted]
- wyldfire 1y agoToo subtle - was this oracle vs java one? Remind me: java won or lost that one?
- mandevil 1y agoOracle sued Google, and Google won, 6-2 (RBG was dead, Barrett had not yet been confirmed when the case was heard). Supreme Court ruled that by applying the Four Factors of Fair Use, Google stayed within Fair Use. An API specification ends up being a system of organizing things, like the Dewey Decimal System (and thus not really something that can be copyrighted), which in the end marks the first factor for Google. Because Google limited the Android version of the API to just things that were useful for smart phones it won on the second factor too. Because only 0.4% of the code was reused, and mostly was rewritten, Google won on the third factor. And on the market factor, if they held for Oracle, it would harm the public because then "Oracle alone would hold the key. The result could well prove highly profitable to Oracle (or other firms holding a copyright in computer interfaces) ... [but] the lock would interfere with, not further, copyright's basic creativity objectives." So therefore the fourth factor was also pointing in Google's favor. Whether "java" won or lost is a question of what is "java"? Android can continue to use the Java API- so it is going to see much more activity. But Oracle didn't get to demand license fees, so they are sad.
- moralestapia 1y agoOh man, thanks for this. I always thought it was resolved as infringement and they had to license the Java APIs or something ... Wow.
- 15155 1y agoConsidering 100% of the low-level CUDA API headers have the word "CUDA" in them, this would be interesting to know.
- mayli 1y agoYeah, nice to have MLX-opencl or MLX-amd-whatever
- almostgotcaught 1y ago> CUDA backend backend
- tekacs 1y agoThis instance is the other way around, but that's what this is – CUDA on AMD (or other platforms): https://docs.scale-lang.com/stable/ https://docs.scale-lang.com/stable/
- pjmlp 1y agoWhy, everyone keeps trying to copy CUDA while failing to understand why many of us love it.
- int_19h 1y agoAbstraction layers for GPU compute already exist; this is yet another one, so it doesn't change anything substantially. Most of the time code written using such layers ends up running on NVIDIA hardware in prod anyway, so if anything that is a net positive for the company - it means that more people can now develop for its hardware on their devices.
- zdw 1y agoHow does this work when one of the key features of MLX is using a unified memory architecture? (see bullets on repo readme: https://github.com/ml-explore/mlx https://github.com/ml-explore/mlx ) I would think that bringing that to all UMA APUs (of any vendor) would be interesting, but discreet GPU's definitely would need a different approach? edit: reading the PR comments, it appears that CUDA supports a UMA API directly, and will transparently copy as needed.
- freeone3000 1y agoEh yes but from my experience its lack of prefetch lends to significant memory stalls waiting for the copy. It might be suitable if your entire dataset fits in VRAM after doing a “manual prefetch” but it killed performance for my application (ML training) so hard that we actually got time to move to streaming loads.
- nerdsniper 1y agoEdit: I had the details of the Google v Oracle case wrong. SCOTUS found that re-implementing an API does not infringe copyright. I was remembering the first and second appellate rulings. Also apparently this is not a re-implementation of CUDA.
- skyde 1y agothis is CUDA backend to MLX not MLX backend for CUDA!
- liuliu 1y agoYou misunderstood and this is not re-implementing CUDA API. MLX is a PyTorch-like framework.
- Uehreka 1y agoThis is exactly the kind of thing I wouldn’t opine on until like, an actual lawyer weighs in after thoroughly researching it. There are just too many shades of meaning in this kind of case law for laymen to draw actionable conclusions directly from the opinions. Though I imagine that if Apple is doing this themselves, they likely know what they’re doing, whatever it is.
- MuffinFlavored 1y agoIs this for Mac's with NVIDIA cards in them or Apple Metal/Apple Silicon speaking CUDA?... I can't really tell. Edit: looks like it's "write once, use everywhere". Write MLX, run it on Linux CUDA, and Apple Silicon/Metal.
- cowsandmilk 1y agoNeither, it is for Linux computers with NVIDIA cards
- MBCook 1y agoSeems you already found the answer. I’ll note Apple hasn’t shipped an Nvidia card in a very very long time. Even on the Mac pros before Apple Silicon they only ever sold AMD cards. My understanding from rumors is that they had a falling out over the problems with the dual GPU MacBook Pros and the quality of drivers. I have no idea if sticking one in on the PCI bus let you use it for AI stuff though.
- kmeisthax 1y agoOn Apple Silicon, writing to memory on a PCIe / Thunderbolt device will generate an exception. ARM spec says you're allowed to write to devices as if they were memory but Apple enforces that all writes to external devices go through a device memory mapping[0]. This makes using an external GPU on Apple Silicon[1] way more of a pain in the ass, if not impossible. AFAIK nobody's managed to write an eGPU driver for Apple Silicon, even with Asahi. [0] https://developer.arm.com/documentation/102376/0200/Device-memory https://developer.arm.com/documentation/102376/0200/Device-m... [1] Raspberry Pi 4's PCIe has the same problem AFAIK
- bobmcnamara 1y agoEwww, that kills out of order CPU performance. If it's like ARMv7, it effectively turns each same-page access into it's own ordering barrier.
- saagarjha 1y ago
- Keyframe 1y agoNow do linux support / drivers for Mac hardware!
- lvl155 1y agoSeriously. Those Apple guys became delusional especially after Jobs passed away. These guys just sat on their successes and did nothing for a decade plus. M1 was nice but that was all Jobs doing and planning. I don’t like this Apple. They forgot how to innovate. But I guess we have a VR device nobody wants.
- jjtheblunt 1y agoIt would be funny if you were typing out your response on an iPhone that has been running for 36 hours without recharging.
- marcellus23 1y ago> M1 was nice but that was all Jobs doing and planning M1 was launched 9 years after Jobs died. You're saying they had everything ready to go back then and just sat on their asses for a decade?
- lvl155 1y agoWho bought Semi? Jobs knew they had to make their own. M1 is just a product of their iPhone chips hence all the efficiency.
- albertzeyer 1y agoThis is exciting. So this is using unified memory of CUDA? I wonder how well that works. Is the behavior of the unified memory in CUDA actually the same as for Apple silicon? For Apple silicon, as I understand, the memory is anyway shared between GPU and CPU. But for CUDA, this is not the case. So when you have some tensor on CPU, how will it end up on GPU then? This needs a copy somehow. Or is this all hidden by CUDA?
- MBCook 1y agoThis is my guess, but does higher end hardware they sell, like the server rack stuff for AI, perhaps have the unified memory? I know standard GPUs don’t. The patch suggested one of the reasons for it was to make it easy to develop on a Mac and run on a super computer. So the hardware with the unified memory might be in that class.
- Y_Y 1y agoThe servers don't, but the Jetsons do
- ajuhasz 1y agoThe Jetsons[1] have unified memory[2]. [1] https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/ https://www.nvidia.com/en-us/autonomous-machines/embedded-sy... [2] https://www.nvidia.com/en-us/on-demand/session/gtcspring22-se2600/ https://www.nvidia.com/en-us/on-demand/session/gtcspring22-s...
- tonyarkles 1y agoThey sure do and it's pretty amazing. One iteration of a vision system I worked on got frames from a camera over a Mellanox NIC that supports RDMA (Rivermax), preprocessed the images using CUDA, did inference on them with TensorRT, and the first time a single byte of the inference pipeline hit the CPU itself was when we were consuming the output.
- patrickkrusiec 1y agoThe physical memory is not be unified, but on modern rack scale Nvidia systems, like Grace Hopper or NVL72, the CPU and the GPU(s) share the same virtual address space and have non-uniform memory access to each other's memory.
- paulirish 1y agoIt's coming from zcbenz who created Electron among others https://zcbenz.com/ https://zcbenz.com/ Nice.
- benreesman 1y agoI wonder how much this is a result of Strix Halo. I had a fairly standard stipend for a work computer that I didn't end up using for a while so I recently cashed it in on the EVO-X2 and fuck me sideways: that thing is easily competitive with the mid-range znver5 EPYC machines I run substitors on. It mops the floor with any mere-mortal EC2 or GCE instance, like maybe some r1337.xxxxlarge.metal.metal or something has an edge, but the z1d.metal and the c6.2xlarge or whatever type stuff (fast cores, good NIC, table stakes), blows them away. And those things are 3-10K a month with heavy provisioned IOPS. This thing has real NVME and it cost 1800. I haven't done much local inference on it, but various YouTubers are starting to call the DGX Spark overkill / overpriced next to Strix Halo. The catch of course is ROCm isn't there yet (they're seeming serious now though, matter of time). Flawless CUDA on Apple gear would make it really tempting in a way that isn't true with Strix so cheap and good.
- jitl 1y agoIt’s pretty explicitly targeting cloud cluster training in the PR description.
- ivape 1y agoIf we believe that there’s not enough hardware to meet demand, then one could argue this helps Apple meet demand, even if it’s just by a few percentage points.
- nl 1y ago> The catch of course is ROCm isn't there yet (they're seeming serious now though, matter of time). Competitive AMD GPU neural compute has been any day now for at least 10 years.
- bigyabai 1y agoThe inference side is fine, nowadays. llama.cpp has had a GPU-agnostic Vulkan backend for a while, it's the training side that tends to be a sticking point for consumer GPUs.
- 1y ago
- orliesaurus 1y agoWhy is this a big deal, can anyone explain if they are familiar with the space?
- elpakal 1y ago> NVIDIA hardware is widely used for academic and massive computations. Being able to write/test code locally on a Mac and then deploy to super computers would make a good developer experience. That one stands out to me as a mac user.
- radicaldreamer 1y agoMacBooks used to use Nvidia GPUs, then Apple had a falling out with Nvidia and the beef stands to this day (Apple didn’t use Nvidia hardware when training it’s own LLMs for Apple Intelligence). I wouldn’t be surprised if within the next few years we see a return of Nvidia hardware to the Mac, probably starting with low volume products like the MacPro, strictly for professional/high-end use cases.
- fooker 1y ago> Apple didn’t use Nvidia hardware when training it’s own LLMs for Apple Intelligence Do you have some links for this?
- almostgotcaught 1y agoPeople on hn make up more BS than your local bar https://www.investors.com/news/technology/apple-stock-apple-joins-ai-data-center-race/ https://www.investors.com/news/technology/apple-stock-apple-...
- tgma 1y agoWhat did the poster make up? There's one line where they speculated about future and a commentary about beef existing to this day which is subjective but the rest of it was 100% factual: Apple relied on Google for training their LLM for various reasons and they did have a beef with NVIDIA re MacBooks a long time ago after which they switched the entire line to AMD Graphics.
- numpad0 1y ago> This PR is an ongoing effort to add a CUDA backend to MLX looks like it allows MLX code to compile and run on x86 + GeForce hardware, not the other way around.
- sciencesama 1y agoApple is planing to build data centers with mseries of chips for both app development, testing and to host external services!
- deleted 1y ago[deleted]
- mr_toad 1y agoIf they were doing that, they wouldn’t need CUDA support. More likely they have internal developers who want to do development on Apple hardware and deploy to Nvidia hardware in production.
- DidYaWipe 1y ago[dead]
- lukev 1y agoSo to make sure I understand, this would mean: 1. Programs built against MLX -> Can take advantage of CUDA-enabled chips but not: 2. CUDA programs -> Can now run on Apple Silicon. Because the #2 would be a copyright violation (specifically with respect to NVidia's famous moat). Is this correct?
- ls612 1y ago#2 would be Google v. Oracle wouldn’t it?
- saagarjha 1y agoNo, it's because doing 2 would be substantially harder.
- lukev 1y agoThere's a massive financial incentive (billions) to allow existing CUDA code to run on non-NVidia hardware. Not saying it's easy, but is implementation difficulty really the blocker?
- natas 1y agothat means the next apple computer is going to use nvidia gpu(s).
- dnchdnd 1y agoRandom aside: A lot of the people working on MLX don't seem to be officially affiliated with Apple at least in a superficial review. See for example: https://x.com/prince_canuma https://x.com/prince_canuma Idly wondering, is Apple bankrolling this but wants to keep it in the DL? There were also rumours the team was looking to move at one point ?
- jpcompartir 1y agoIt seems more like Open Source devs who are looking to build clout/rep with MLX? Pretty sure Claude Sonnet is actually doing most of the work.
- Abishek_Muthian 1y agoI’ve been very impressed with MLX models; I can open up local models to everyone in the house, something I wouldn’t dare with my Nvidia computer for the risk of burning down the house. I’ve been hoping Apple Silicon becomes a serious contender for Nvidia chips; I wonder if the CUDA support is just Embrace, extend, and extinguish (EEE).
- m3kw9 1y agoI thought you either use MLX for apple silicone or you compile it for cudaw
- neurostimulant 1y ago> Being able to write/test code locally on a Mac and then deploy to super computers would make a good developer experience. Does this means you can use MLX on linux now? Edit: Just tested it and it's working but only python 3.12 version is available on pypi right now: https://pypi.org/project/mlx-cuda/#files https://pypi.org/project/mlx-cuda/#files
- qwertox 1y agoIf Apple would support Nvidia cards it would be the #1 solution for developers.
- Nevermark 1y agoIf Apple doubled the specs of their Ultra M processor every year, in numbers of cores, RAM cells, internal and external bandwidth, until both the Ultra processor and its RAM plane took up full wafers, .... but still fit in a Mac Studio case, with a new white reverse-power heat extraction USB-C+ cable designed to be terminated at a port on a small wireless heat exchanger dish, which instantly beamed all the waste heat into space, at such high efficiency that the Studio internals could operate at -100 Celsius, and all those cores overclocked, oh they over clocked, ... Yes we can dream! It would great if Apple continues pushing M processors to next levels, in part, to go vertical into the cloud. Or if they start supporting nVidia. The latter seems less Apple-y. But they must be considering the value of a cloud level Apple-friendly AI computing solution, so something is likely (?) to happen.
- bigyabai 1y agoNvidia already maintains BSD-native drivers that macOS can use. The only work Apple has to do is implement EGLStream in Quartz, which is probably already done. And even that isn't necessary to get CUDA or video acceleration working. In theory, the only thing stopping Nvidia from putting those features on macOS is signed driver code. It's not rocket science, nobody has to "dream" for such an ordinary feature. On Linux and Windows it's treated as table-stakes, which is part of why they're desirable cloud OSes and macOS isn't. Apple certainly knows as much, it's not like they have customers begging to bring back Xserve.
- adultSwim 1y agoThis is great to see. I had wrongly assumed MLX was Apple-only.
- neuroelectron 1y agoJust remember to name for fp8 kernels "cutlass" for +50% performance.
- mattfrommars 1y agoIt's year 2025 and we have yet to have impact of CUDA like what Java had in the idea, "write once, run it anywhere" Academia and companies continue to write proprietary code. Its as if we continue to write code for Adobe Flash or Microsoft Silverlight in year 2025. Honestly, I don't mind as Nvidia shareholder.
- raincole 1y agoIn the end Java doesn't achieve "write once, run it anywhere" either. I guess there might be a way to develop apps for iOS or even PlayStation in Java, but my knees hurt just thinking about how many hoops one needs to jump through.
- bshacklett 1y agoI’m currently working on migrating Java code from mid-range systems to containers in the cloud. The number of code changes required is near zero. They may not have gotten portability perfectly solved, but it’s pretty darn good compared to many other platforms. Now, if the industry could just get out of the ridiculous Java 8/11 rut, we’d be in good shape.
- bigyabai 1y agoI'll never get over the way Apple treated OpenCL. They saw the train coming down the tracks, spent so long hedging their bet against CUDA, and threw in the towel the moment actual demand started cropping up. CUDA very nearly had a serious, corporate-funded and write-once-run-anywhere competitor. Normally I write something snide about not seeing where the puck was headed. But Apple did skate to the puck the puck here, they just did nothing with it.
- int_19h 1y agoBack in the day, the reason why people kept targeting Flash is because all the other alternatives were worse. If you recall, the only thing that made a difference was mobile, where Flash ended up being a liability due to performance and battery lifetime issues. And even then it took a company like Apple, which could rely on its cult status to draw the hard line on mobile Flash, ship iPhone without it (unlike Android which had it, warts and all), and steadfastly refuse to even consider adding it, forcing everybody else to choose between using Flash and supporting the lucrative iPhone ecosystem. I'm not even sure what the equivalent would be for CUDA tbh.