31 ms·
macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt
- nodesocket 10mo agoCan we get proper HDR support first in macOS? If I enable HDR on my LG OLED monitor it looks completely washed out and blacks are grey. Windows 11 HDR works fine.
- Razengan 10mo agoReally? I thought it's always been that HDR was notorious on Windows, hopeless on Linux, and only really worked in a plug-and-play manner on Mac, unless your display has an incorrect profile or something/ https://www.youtube.com/shorts/sx9TUNv80RE https://www.youtube.com/shorts/sx9TUNv80RE
- heavyset_go 10mo agoWorks well on Linux, just toggle a checkmark in the settings.
- masspro 10mo agoMacOS does wash out SDR content in HDR mode specifically on non-Apple monitors. An HDR video playing in windowed mode will look fine but all the UI around it has black and white levels very close to grey. Edit: to be clear, macOS itself (Cocoa elements) is all SDR content and thus washed out.
- Starmina 10mo agoThat's intended behavior for monitor limited in peak brightness
- nodesocket 10mo agoI don't think so. Windows 11 has a HDR calibration utility that allows you to adjust brightness and HDR and it maintains blacks being perfectly black (especially with my OLED). When I enable HDR on macOS whatever settings I try, including adjusting brightness and contrast on the monitor the blacks look completely washed out and grey. HDR DOES seem to work correctly on macOS but only if you use Mac displays.
- masspro 10mo agoThat’s the statement I found last time I went down this rabbit hole, that they don’t have physical brightness info for third-party displays so it just can’t be done any better. But I don’t understand how this can lead to making the black point terrible. Black should be the one color every emissive colorspace agrees on.
- kmeisthax 10mo agoActually, intended behavior in general. Even on their own displays the UI looks grey when HDR is playing. Which, personally, I find to be extremely ugly and gross and I do not understand why they thought this was a good idea.
- adastra22 10mo agoHuh, so that’s why HDR looks like shit on my Mac Studio.
- crazygringo 10mo agoDefine "washed out"? The white and black levels of the UX are supposed to stay in SDR. That's a feature not a bug. If you mean the interface isn't bright enough, that's intended behavior. If the black point is somehow raised, then that's bizarre and definitely unintended behavior. And I honestly can't even imagine what could be causing that to happen. It does seem like that it would have to be a serious macOS bug. You should post a photo of your monitor, comparing a black #000 image in Preview with a pitch-black frame from a video. People edit HDR video on Macs, and I've never heard of this happening before.
- robflynn 10mo agoOh, that explains why it looked so odd when I enabled HDR on my Studio.
- m-ack-toddler 10mo agoAI is arguably more important than whatever gaming gimmick you're talking about.
- djdkdldl 10mo ago[flagged]
- simonw 10mo agoI follow the MLX team on Twitter and they sometimes post about using MLX on two or more joined together Macs to run models that need more than 512GB of RAM. A couple of examples: Kimi K2 Thinking (1 trillion parameters): https://x.com/awnihannun/status/1986601104130646266 https://x.com/awnihannun/status/1986601104130646266 DeepSeek R1 (671B): https://x.com/awnihannun/status/1881915166922863045 https://x.com/awnihannun/status/1881915166922863045 - that one came with setup instructions in a Gist: https://gist.github.com/awni/ec071fd27940698edd14a4191855bba6 https://gist.github.com/awni/ec071fd27940698edd14a4191855bba...
- awnihannun 10mo agoFor a bit more context, those posts are using pipeline parallelism. For N machines put the first L/N layers on machine 1, next L/N layers on machine 2, etc. With pipeline parallelism you don't get a speedup over one machine - it just buys you the ability to use larger models than you can fit on a single machine. The release in Tahoe 26.2 will enable us to do fast tensor parallelism in MLX. Each layer of the model is sharded across all machines. With this type of parallelism you can get close to N-times faster for N machines. The main challenge is latency since you have to do much more frequent communication.
- liuliu 10mo agoBut that's only for prefilling right? Or is it beneficial for decoding too (I guess you can do KV lookup on shards, not sure how much speed-up that will be though).
- zackangelo 10mo agoNo you use tensor parallelism in both cases. The way it typically works in an attention block is: smaller portions of the Q, K and V linear layers are assigned to each node and are processed independently. Attention, rope norm etc is run on the node-specific output of that. Then, when the output linear layer is applied an "all reduce" is computed which combines the output of all the nodes. EDIT: just realized it wasn't clear -- this means that each node ends up holding a portion of the KV cache specific to its KV tensor shards. This can change based on the specific style of attention (e.g., in GQA where there are fewer KV heads than ranks you end up having to do some replication etc)
- pstuart 10mo agoI imagine that M5 Ultra with Thunderbolt 5 could be a decent contender for building plug and play AI clusters. Not cheap, but neither is Nvidia.
- whimsicalism 10mo agonvidia is absolutely cheaper per flop
- FlacksonFive 10mo agoTo acquire, maybe, but to power?
- whimsicalism 10mo agomachine capex currently dominates power
- amazingman 10mo agoSounds like an ecosystem ripe for horizontally scaling cheaper hardware.
- crote 10mo agoIf I understand correctly, a big problem is that the calculation isn't embarrasingly parallel: the various chunks are not independent, so you need to do a lot of IO to get the results from step N from your neighbours to calculate step N+1. Using more smaller nodes means your cross-node IO is going to explode. You might save money on your compute hardware, but I wouldn't be surprised if you'd end up with an even greater cost increase on the network hardware side.
- adastra22 10mo agoFLOPS are not what matters here.
- jeffbee 10mo agoVery cool. It requires a fully-connected mesh so the scaling limit here would seem to be 6 Mac Studio M3 Ultra, up to 3TB of unified memory to work with.
- PunchyHamster 10mo agoI'm sure someone will figure out how to make thunderbolt switch/router
- huslage 10mo agoI don't believe the standard supports such a thing. But I wonder if TB6 will.
- kmeisthax 10mo agoRDMA is a networking standard, it's supposed to be switched. The reason why it's being done over Thunderbolt is that it's the only cheap/prosumer I/O standard with enough bandwidth to make this work. Like, 100Gbit Ethernet cards are several hundred dollars minimum, for two ports, and you have to deal with SFP+ cabling. Thunderbolt is just way nicer[0]. The way this capability is exposed in the OS is that the computers negotiate an Ethernet bridge on top of the TB link. I suspect they're actually exposing PCIe Ethernet NICs to each other, but I'm not sure. But either way, a "Thunderbolt router" would just be a computer with a shitton of USB-C ports (in the same way that an "Ethernet router" is just a computer with a shitton of Ethernet ports). I suspect the biggest hurdle would actually just be sourcing an SoC with a lot of switching fabric but not a lot of compute. Like, you'd need Threadripper levels of connectivity but with like, one or two actual CPU cores. [0] Like, last time I had to swap work laptops, I just plugged a TB cable between them and did an `rsync`.
- bleepblap 10mo agoI think you might be swapping RDMA with RoCE - RDMA can happen entirely within a single node. For example between an NVME and a GPU.
- novok 10mo agoNow we need some hardware that is rackmount friendly, an OS that is not fidly as hell to manage in a data center or headless server and we are off to the races! And no, custom racks are not 'rackmount friendly'.
- joeframbach 10mo agoSo, the Powerbook Duo Dock?
- btown 10mo agoIt would be incredibly ironic if, with Apple's relatively stable supply chain relative to the chaos of the RAM market these days (projected to last for years), Apple compute became known as a cost-effective way to build medium-sized clusters for inference.
- andy99 10mo agoIt’s gonna suck if all the good Macs get gobbled up by commercial users.
- mschuster91 10mo agoit's not like regular people can afford this kind of Apple machine anyway.
- teeray 10mo agoIt’s just depressing that the “PC in every home” era is being rapidly pulled out from under our feet by all these supply shocks.
- dghlsakjg 10mo agoHuh? Home PCs are as cheap as they’ve ever been. Adjusted for inflation the same can be said about “home use” Macs. The list price of an entry level MacBook Air has been pretty much the same for more than a decade. Adjust for inflation, and you get a MacBook air for less than half the real cost of the launch model that is massively better in every way. A blip in high end RAM prices has no bearing on affordable home computing. Look at the last year or two and the proliferation of cheap, moderately to highly speced mini desktops. I can get a Ryzen 7 system with 32gb of ddr5, and a 1tb drive delivered to my house before dinner tomorrow for $500 + tax. That’s not depressing, that’s amazing!
- behnamoh 10mo ago> Home PCs are as cheap as they’ve ever been. just the 5090 GPU costs +$3k, what are you even talking about
- timsneath 10mo agoAlso see https://www.engadget.com/ai/you-can-turn-a-cluster-of-macs-into-an-ai-supercomputer-in-macos-tahoe-262-191500778.html https://www.engadget.com/ai/you-can-turn-a-cluster-of-macs-i...
- geerlingguy 10mo agoThis implies you'd run more than one Mac Studio in a cluster, and I have a few concerns regarding Mac clustering (as someone who's managed a number of tiny clusters, with various hardware): 1. The power button is in an awkward location, meaning rackmounting them (either 10" or 19" rack) is a bit cumbersome (at best) 2. Thunderbolt is great for peripherals, but as a semi-permanent interconnect, I have worries over the port's physical stability... wish they made a Mac with QSFP :) 3. Cabling will be important, as I've had tons of issues with TB4 and TB5 devices with anything but the most expensive Cable Matters and Apple cables I've tested (and even then...) 4. macOS remote management is not nearly as efficient as Linux, at least if you're using open source / built-in tooling To that last point, I've been trying to figure out a way to, for example, upgrade to macOS 26.2 from 26.1 remotely, without a GUI, but it looks like you _have_ to use something like Screen Sharing or an IP KVM to log into the UI, to click the right buttons to initiate the upgrade. Trying "sudo softwareupdate -i -a" will install minor updates, but not full OS upgrades, at least AFAICT.
- eurleif 10mo agoI have no experience with this, but for what it's worth, looks like there's a rack mounting enclosure available which mechanically extends the power switch: https://www.sonnetstore.com/products/rackmac-studio https://www.sonnetstore.com/products/rackmac-studio
- geerlingguy 10mo agoI have something similar from MyElectronics, and it works, but it's a bit expensive, and still imprecise. At least the power button isn't in the back corner underneath!
- wlesieutre 10mo agoFor #2, OWC puts a screw hole above their dock's thunderbolt ports so that you can attach a stabilizer around the cord https://www.owc.com/solutions/thunderbolt-dock https://www.owc.com/solutions/thunderbolt-dock It's a poor imitation of old ports that had screws on the cables, but should help reduce inadvertent port stress. The screw only works with limited devices (ie not the Mac Studio end of the cord) but it can also be adhesive mounted. https://eshop.macsales.com/item/OWC/CLINGON1PK/ https://eshop.macsales.com/item/OWC/CLINGON1PK/
- deleted 10mo ago[deleted]
- givemeethekeys 10mo agoWould this also work for gaming?
- AndroTux 10mo agoNo
- storus 10mo agoIs there any way to connect DGX Sparks to this via USB4? Right now only 10GbE can be used despite both Spark and MacStudio having vastly faster options.
- zackangelo 10mo agoSparks are built for this and actually have Connect-X 7 NICs built in! You just need to get the SFPs for them. This means you can natively cluster them at 200Gbps.
- wtallis 10mo agoThat doesn't answer the question, which was how to get a high-speed interconnect between a Mac and a DGX Spark. The most likely solution would be a Thunderbolt PCIe enclosure and a 100Gb+ NIC, and passive DAC cables. The tricky part would be macOS drivers for said NIC.
- zackangelo 10mo agoYou’re right I misunderstood. I’m not sure if it would be of much utility because this would presumably be for tensor parallel workloads. In that case you want the ranks in your cluster to be uniform or else everything will be forced to run at the speed of the slowest rank. You could run pipeline parallel but not sure it’d be that much better than what we already have.
- storus 10mo agoIt was about this use case: https://blog.exolabs.net/nvidia-dgx-spark/ https://blog.exolabs.net/nvidia-dgx-spark/
- daft_pink 10mo agoHoping Apple has secured plentiful DDR5 to use in their machines so we can buy M5 chips with massive amounts of RAM soon.
- colechristensen 10mo agoApple tends to book its fab time / supplier capacity years in advance
- lossolo 10mo agoI hope so, I want to replace my M1 Pro with MacBook Pro with M5 Pro when they release it next year.
- colechristensen 10mo agoI mostly want the M5 Pro because my choice of an M4 Air this year with 24 GB of RAM is turning out to be less than I want with the things I'm doing these days.
- reaperducer 10mo agoAs someone not involved in this space at all, is this similar to the old MacOS Xgrid? https://en.wikipedia.org/wiki/Xgrid https://en.wikipedia.org/wiki/Xgrid
- wmf 10mo agoNo.
- reilly3000 10mo agodang I wish I could share md tables. Here’s a text edition: For $50k the inference hardware market forces a trade-off between capacity and throughput: * Apple M3 Ultra Cluster ($50k): Maximizes capacity (3TB). It is the only option in this price class capable of running 3T+ parameter models (e.g., Kimi k2), albeit at low speeds (~15 t/s). * NVIDIA RTX 6000 Workstation ($50k): Maximizes throughput (>80 t/s). It is superior for training and inference but is hard-capped at 384GB VRAM, restricting model size to <400B parameters. To achieve both high capacity (3TB) and high throughput (>100 t/s) requires a ~$270,000 NVIDIA GH200 cluster and data center infrastructure. The Apple cluster provides 87% of that capacity for 18% of the cost.
- mechagodzilla 10mo agoYou can keep scaling down! I spent $2k on an old dual-socket xeon workstation with 768GB of RAM - I can run Deepseek-R1 at ~1-2 tokens/sec.
- ternus 10mo agoAnd if you get bored of that, you can flip the RAM for more than you spent on the whole system!
- Weryj 10mo agoJust keep going! 2TB of swap disk for 0.0000001 t/sec
- kergonath 10mo agoHang on, starting benchmarks on my Raspberry Pi.
- euroderf 10mo agoBy the year 2035, toasters will run LLMs.
- pickle-wizard 10mo ago
- ComputerGuru 10mo agoImagine if the Xserve was never killed off. Discontinued 14 years ago, now!
- icedchai 10mo agoIf it was still around, it would probably still be stuck on M2, just like the Mac Pro.
- stego-tech 10mo agoThis doesn’t remotely surprise me, and I can guess Apple’s AI endgame: * They already cleared the first hurdle to adoption by shoving inference accelerators into their chip designs by default. It’s why Apple is so far ahead of their peers in local device AI compute, and will be for some time. * I suspect this introduction isn’t just for large clusters, but also a testing ground of sorts to see where the bottlenecks lie for distributed inference in practice. * Depending on the telemetry they get back from OSes using this feature, my suspicion is they’ll deploy some form of distributed local AI inference system that leverages their devices tied to a given iCloud account or on the LAN to perform inference against larger models, but without bogging down any individual device (or at least the primary device in use) For the endgame, I’m picturing a dynamically sharded model across local devices that shifts how much of the model is loaded on any given device depending on utilization, essentially creating local-only inferencing for privacy and security of their end users. Throw the same engines into, say, HomePods or AppleTVs, or even a local AI box, and voila, you’re golden. EDIT: If you're thinking, "but big models need the higher latency of Thunderbolt" or "you can't do that over Wi-Fi for such huge models", you're thinking too narrowly. Think about the devices Apple consumers own, their interconnectedness, and the underutilized but standardized hardware within them with predictable OSes. Suddenly you're not jamming existing models onto substandard hardware or networks, but rethinking how to run models effectively over consumer distributed compute. Different set of problems.
- threecheese 10mo agoI think you are spot on, and this fits perfectly within my mental model of HomeKit; tasks are distributed to various devices within the network based on capabilities and authentication, and given a very fast bus Apple can scale the heck out of this.
- stego-tech 10mo agoConsumers generally have far more compute than they think; it's just all distributed across devices and hard to utilize effectively over unreliable interfaces (e.g. Wi-Fi). If Apple (or anyone, really) could figure out a way to utilize that at modern scales, I wager privacy-conscious consumers would gladly trade some latency in responses in favor of superior overall model performance - heck, branding it as "deep thinking" might even pull more customers in via marketing alone ("thinks longer, for better results" or some vaguely-not-suable marketing slogan). It could even be made into an API for things like batch image or video rendering, but without the hassle of setting up an app-specific render farm. There's definitely something there, but Apple's really the only player setup to capitalize on it via their halo effect with devices and operating systems. Everyone else is too fragmented to make it happen.
- 650REDHAIR 10mo agoDo we think TB4 is on the table or is there a technical limitation?
- piskov 10mo agoGeorge Hotz made nvidia running on macs with his tinygrad via usb4 https://x.com/__tinygrad__/status/1980082660920918045 https://x.com/__tinygrad__/status/1980082660920918045
- throawayonthe 10mo agohttps://social.treehouse.systems/@janne/115509948515319437 https://social.treehouse.systems/@janne/115509948515319437 nvidia on a 2023 Mac Pro running linux :p
- piskov 10mo agoGeohotz stuff anyone can run today
- londons_explore 10mo agoNobodies gonna take them seriously till they make something rack mounted and that isn't made of titanium with pentalobe screws...
- moralestapia 10mo agoYou might ignore this but, for a while, Mac Mini clusters were a thing and they were capex and opex effective. That same setup is kind of making a comeback.
- londons_explore 10mo agoIt's in a similar vein to the PS2 linux cluster or someone trying to use vape CPU's as web servers... It might be cost effective, but the supplier is still saying "you get no support, and in fact we might even put roadblocks in your way because you aren't the target customer".
- moralestapia 10mo agoTrue. I'm sure Apple could make a killing on the server side, unfortunately their income from their other products is so big that even if that's a 10B/year opportunity they'll be like "yawn, yeah, whatever".
- fennecbutt 10mo agoDoubt. A 10B idea is still a promotion. And if capitalism is shrinkflationing hard, which it is atm, then capitalists would not leave something like that on the table.
- fennecbutt 10mo agoThey were only a thing to do ci/compilation related to apples os because their walled garden locked using other platforms out. You're building an iPhone or mac app? Well your ci needs to be on a cluster of apple machines.
- cluckindan 10mo agoThis sounds like a plug’n’play physical attack vector.
- guiand 10mo agoFor security, the feature requires setting a special option with the recovery mode command line: rdma_ctl enable
- int32_64 10mo agoApple should setup their own giant cloud of M chips with tons of vram, make Metal as good as possible for AI purposes, then market the cloud as allowing self-hosted models for companies and individuals that care about privacy. They would clean up in all kinds of sectors whose data can't touch the big LLM companies.
- wmf 10mo agoThat exists but it's only for iUsers running Apple models. https://security.apple.com/blog/private-cloud-compute/ https://security.apple.com/blog/private-cloud-compute/
- make3 10mo agoThe advantages of having a single big memory per gpu are not as big in a data center where you can just shard things between machines and use the very fast interconnect, saturating the much faster compute cores of a non Apple GPU from Nvidia or AMD
- sebnukem2 10mo agoI didn't know they skipped 10 version numbers.
- badc0ffee 10mo agoThey switched to using the year.
- thatwasunusual 10mo agoCan someone do an ELI5, and why this is important?
- wmf 10mo agoIt's faster and lower latency than standard Thunderbolt networking. Low latency makes AI clusters faster.
- schmuckonwheels 10mo agoThat's nice but Liquid (gl)ass still sucks.
- yalogin 10mo agoAs someone that is not familiar with rdma, dos it mean I can connect multiple Macs and run inference? If so it’s great!
- wmf 10mo agoYou've been able to run inference on multiple Macs for around a year but now it's much faster.
- 0manrho 10mo agoJust for reference: Thunderbolt5's stated "80Gbps" bandwidth comes with some caveats. That's the figure for either Display Port bandwidth itself or in practice more often realized by combining the data channel (PCIe4x4 ~=64Gbps) with the display channels (=<80Gbps if used in concert with data channels), and potentially it can also do unidirectional 120Gbps of data for some display output scenarios. If Apple's silicon follows spec, then that means you're most likely limited to PCIe4x4 ~=64Gbps bandwidth per TB port, with a slight latency hit due to the controller. That Latency hit is ItDepends(TM), but if not using any other IO on that controller/cable (such as display port), it's likely to be less than 15% overhead vs Native on average, but depending on drivers, firmware, configuration, usecase, cable length, and how apple implemented TB5, etc, exact figures very. And just like how 60FPS Average doesn't mean every frame is exactly 1/60th of a second long, it's entirely possible that individual packets or niche scenarios could see significantly more latency/overhead. As a point of reference Nvidia RTX Pro (formerly known as quadro) workstation cards of Ada generation and older along with most modern consumer grahics cards are PCIe4 (or less, depending on how old we're talking), and the new RTX Pro Blackwell cards are PCIe5. Though comparing a Mac Studio M4 Max for example to an Nvidia GPU is akin to comparing Apples to Green Oranges However, I mention the GPU's not just to recognize the 800lb AI compute gorilla in the room, but also that while it's possible to pool a pair of 24GB VRAM GPU's to achieve a 48GB VRAM pool between them (be it through a shared PCIe bus or over NVlink), the performance does not scale linearly due to PCIe/NVLinks limitations, to say nothing of the software, and configuration and optimization side of things also being a challenge to realizing max throughput in practice. This is also just as true as a pair of TB5 equipped macs with 128GB of memory each using TB5 to achieve a 256GB Pool will take a substantial performance hit compared to on otherwise equivalent mac with 256GB. (capacities chosen are arbitrary to illustrate the point). The exact penalty really depends on usecase and how sensitive it is to the latency overhead of using TB5 as well as the bandwidth limitation. It's also worth noting that it's not just entirely possible with RDMA solutions (no matter the specifics) to see worse performance than using a singular machine if you haven't properly optimized and configured things. This is not hating on the technology, but a warning from experience for people who may have never dabbled to not expect things to just "2x" or even just better than 1x performance just by simply stringing a cable between two devices. All that said, glad to see this from Apple. Long overdue in my opinion as I doubt we'll see them implement an optical network port with anywhere near that bandwidth or RoCEv2 support, much less a expose a native (not via TB) PCIe port on anything that's a non-pro model. EDIT: Note, many mac skus have multiple TB5 ports, but it's unclear to me what the underlying architecture/topology is there and thus can't speculate on what kind of overhead or total capacity any given device supports by attempting to use multiple TB links for more bandwidth/parallelism. If anyone's got an SoC diagram or similar refernce data that actually tells us how the TB controller(s) are uplinked to the rest of the SoC, I could go in more depth there. I'm not an Apple silicon/MacOS expert. I do however have lots of experience with RDMA/RoCE/IB clusters, NVMeoF deployments, SXM/NVlink'd devices and generally engineering low latency/high performance network fabrics for distributed compute and storage (primarily on the infrastructure/hardware/ops side than on the software side) so this is my general wheelhouse, but Apple has been a relatively blindspot for me due to their ecosystem generally lacking features/support for things like this.
- kjkjadksj 10mo agoRemember when they enabled egpu over thunderbolt and no one cared because the thunderbolt housing cost almost as much as your macbook outright? Yeah. Thunderbolt is a racket. It’s a god damned cord. Why is it $50.
- wmf 10mo agoIn this case Thunderbolt is much much cheaper than 100G Ethernet. (The cord is $50 because it contains two active chips BTW.)
- geerlingguy 10mo agoYeah, even decent 40 Gbps QSFP+ DAC cables are usually $30+, and those don't have active electronics in them like Thunderbolt does. The ability to also deliver 240W (IIRC?) over the same cable is also a bit different here, it's more like FireWire than a standard networking cable.
- FridgeSeal 10mo agoThat’s great for AI people, but can we use this for other distributed workloads that aren’t ML?
- geerlingguy 10mo agoI've been testing HPL and mpirun a little, not yet with this new RDMA capability (it seems like Ring is currently the supported method)... but it was a little rough around the edges. See: https://ml-explore.github.io/mlx/build/html/usage/distributed.html#getting-started-with-mpi https://ml-explore.github.io/mlx/build/html/usage/distribute...
- dagmx 10mo agoSure, there’s nothing about it that’s tied to ML. It’s faster interconnect , use it for many kinds of shared compute scenarios.
- nickysielicki 10mo agoThis is such a weird project. Like where is this running at scale? Where’s the realistic plan to ever run this at scale? What’s the end goal here? Don’t get me wrong... It’s super cool, but I fail to understand why money is being spent on this.
- aurareturn 10mo agoThe end goal is that Macs become good local LLM inference machines and for AI devs to keep using Macs.
- nickysielicki 10mo agoThe former will never happen and the latter is a certainty.
- aurareturn 10mo agoThe former is already true and will become even more true when M5 Pro/Max/Ultra release.
- pjmlp 10mo agoMaybe Apple should rethink bringing back Mac Pro desktops with pluggable GPUs, like that one in the corner still playing with its Intel and AMD toys, instead of a big box full of air and pro audio cards only.
- nottorp 10mo agoIt's good to sell shovels :)
- zeristor 10mo agoWill Apple be able to ramp up M3 Ultra MacStudios if this becomes a big thing? Is this part of Apple’s plan of building out server side AI support using their own hardware? If so they would need more physical data centres. I’m guessing they too would be constrained by RAM.
- unit149 10mo agoGarageband DAW + MacOS 14.4 Roland Juno-D7 synthsizer, for 8-bit audio complementary compact disk format as AIFF, WAV, or MIDI appliance, in which under SLA-royalties licenses, binary 44.1 Khz sample rate sets the reproducer for reference level. [1]: https://www.apple.com/legal/sla/docs/GarageBand.pdf https://www.apple.com/legal/sla/docs/GarageBand.pdf
- irusensei 10mo agoI am waiting for M5 studio but due to current price of hardware I'm not sure it will be at a level that I would call affordable. Currently I'm watching for news and if there is any announcement prices will go up I'll probably settle for an M4 Max.
- DesiLurker 10mo agodoes this means an egpu might finally work with macbook-pro or studio?
- wmf 10mo agoNo.
- jamesfmilne 10mo agoAnyone found any APIs related to this? I'd have some other uses for RDMA between Macs.
- jamesfmilne 10mo agoI found some useful clues here. Looks like it uses the regular InfiniBand RDMA APIs. https://github.com/Anemll/mlx-rdma/commit/a901dbd3f9eeefc62800da1142702bda8ea4f637 https://github.com/Anemll/mlx-rdma/commit/a901dbd3f9eeefc628...
- TheRealPomax 10mo agoIS this... good? Why is this something that the underlying OS itself should be involved in at all?
- wmf 10mo agoNetworking is part of the OS's job.
- sora2video 10mo ago[dead]