11 ms·
512GB of unified memory is truly breaking new ground. I was wondering when Apple would overcome memory constraints, and now we're seeing a half-terabyte level o
by cxie 2y ago
512GB of unified memory is truly breaking new ground. I was wondering when Apple would overcome memory constraints, and now we're seeing a half-terabyte level of unified memory. This is incredibly practical for running large AI models locally ("600 billion parameters"), and Apple's approach of integrating this much efficient memory on a single chip is fascinating compared to NVIDIA's solutions.
I'm curious about how this design of "fusing" two M3 Max chips performs in terms of heat dissipation and power consumption though
- samstave 2y ago"unified memory" funny that people think this is so new, when CRAY had Global Heap eons ago...
- ddtaylor 2y agoWhy did it take so long for us to get here?
- baby_souffle 2y agoJust a guess, but fabricating this can't be easy. Yield is probably higher if you have less memory per chip.
- astrange 2y agoIt's regular memory on separate chips.
- RachelF 2y agoSome possible groups of reasons: 1. Until recently RAM amount was something the end user liked to configure, so little market demand. 2. Technically, building such a large system on a chip or collection of chiplets was not possible. 3. RAM speed wasn't a bottleneck for most tasks, it was IO or CPU. LLMs changed this.
- hot_gril 2y agoM1 came out before the LLM rush, though
- philistine 2y agoApple has always liked to integrate as much as possible on the same chip. It was only natural that they would come to this conclusion, with the improved perf the cherry on top.
- hot_gril 2y agoWell also these chips originated in phones, where they kinda had to integrate it. And the quicker RAM and disk access are pretty nice.
- wtallis 2y agoThe M1 is in a product segment where discrete GPUs have been gone for decades, in favor of integrated graphics that shares one pool of RAM with the CPU. The better question to ask is why Apple kept using that unified memory design even when moving up to larger chips like the M1 Max and M1 Ultra.
- MBCook 2y agoThe GPU is built into the same physical die as the CPU. So if you wanted to give it a second ram pool you would have to add an entire second memory interface just for the on-die GPU. Now all you’ve done is make it more complicated, slower because now you have to move things between the two pools, and gained what exactly? I think it was a very clear and obvious decision to make. It’s an outgrowth out of how the base chips were designed, and it turned out to be extremely handy for some things. Plus since all their modern devices now work this way that probably simplify the software. I’m not saying it’s genius foresight, but it certainly worked out rather well. There’s nothing stopping them from supporting discreet GPUs too if they wanted to. They just clearly don’t.
- RachelF 2y ago
- wmf 2y agoLaptops have had unified memory for ten years or more. For desktops very few apps benefit from unified memory.
- djmips 2y agoAnd game consoles that use similar parts as laptops.
- webworker 2y agoThe real hardware needed for artificial intelligence wasn't NVIDIA, it was a CRAY XMP from 1982 all along
- samstave 2y agoWHen I was with Mirantis, I flew to Austin TX to meet a client in a non-descript multi-tenant office building... we walked in and getting our bearings, we come upon CRAY office. WTF?! I tried the doors, locked - and it was clearly empty... but damn did I want to steal their office door signage.
- hot_gril 2y agoIt's new for mainstream PCs to have it.
- TylerE 2y agoNew for performance machines maybe. I remember "integrated graphics" when that meant some shitty co-processor and 16 or 32MB of semi-reserved system RAM.
- pjmlp 2y agoNope, it was common in 8 and 16 bit home computers, and in respect to PCs themselves graphics memory was mapped into the main memory until the arrival of 3D dedicated cards. And even with 3D, integrated GPUs have existed for years.
- djmips 2y agoLike pretty much every game console.
- hot_gril 2y agoThe CPUs with iGPUs didn't also have the memory on-chip. The Nintendo 64 did. Not sure about the old home computers, but I thought those had separate memory usually.
- pjmlp 2y agoOf course not, because they are not designed as SOCs, the only memory on chip is cache, it doesn't change the fact the memory is one whole block shared between CPU and iGPU.
- angoragoats 2y agoApple does not have the memory on-chip (on the same die as the CPU) either.
- Vilian 2y agoIt's not new for PC to block user ram upgrade
- grandempire 2y agoYou mean the room sized super computer than sold tens of units?
- samstave 2y agoYes, but now its in my pocket.
- ProAm 2y agoThis is just Apple disrespecting their customer base.
- bigyabai 2y agoFor enterprise markets, this is table stakes. A lot of datacenter customers will probably ignore this release altogether since there isn't a high-bandwidth option for systems interconnect.
- pavlov 2y agoThe Mac Studio isn’t meant for data centers anyway? It’s a small and silent desktop form factor — in every respect the opposite of a design you’d want to put in a rack. A long time ago Apple had a rackmount server called Xserve, but there’s no sign that they’re interested in updating that for the AI age.
- bigyabai 2y agoIt's the Ultra chip, the same one that goes into the rackmount Mac Pro. I don't think there's much confusion as to who this is for. > there’s no sign that they’re interested in updating that for the AI age. https://security.apple.com/blog/private-cloud-compute/ https://security.apple.com/blog/private-cloud-compute/
- pavlov 2y agoI genuinely forgot the Mac Pro still exists. It’s been so long since I even saw one. And I’ve had every previous Mac tower design since 1999: G4, G5, the excellent dual Xeon, the horrible black trash can… But Apple Silicon delivers so much punch in the Studio form factor, the old school Pro has become very niche. Edit - looks like the new M3 Ultra is only available in Mac Studio anyway? So the existence of the Pro is moot here.
- choilive 2y agonever understood the hate on the trash can. Isn't the mac studio basically the same idea as the trash can but even less upgradeable?
- 2y ago
- FloatArtifact 2y agoThey didn't increase the memory bandwidth. You can get the same memory bandwidth, which is available on the M2 Studio. Yes, yes, of course you can get 512 gigabytes of uRAM for 10 grand. The the question is if a llm will run with usable performance at that scale? The point is there's diminishing returns despite having enough uRAM with the same amount of memory bandwidth even with increased processing speed of the new chip for AI. So there must be a min-max performance ratio between memory bandwidth and the size of the memory pool in relation to the processing power.
- cxie 2y agoGuess what? I'm on a mission to completely max out all 512GB of mem...maybe by running DeepSeek on it. Pure greed!
- gustomksimus25 2y ago[dead]
- swivelmaster 2y agoYou could always just open a few Chrome tabs…
- DidYaWipe 2y ago[flagged]
- ksec 2y ago>Edit: WTF, someone downvoted "Enjoy the upvotes?" Pathetic. You should read HN posting Guidelines if you want to understand why. Although I guess mostly in this case it is someone fat thumbed downvote.
- School-Cotton 2y agoI downvote all Reddit-style memes, jokes, reference humor, catchphrases, and so on. It’s low-effort content that doesn’t fit the vibe of HN and actively makes the site worse for its intended purpose.
- dheera 2y agoIt will cost 4X what it costs to get 512GB on an x86 server motherboard.
- smith7018 2y agoYou can build an x86 machine that can fully run DeepSeek R1 with 512GB VRAM for ~$2,500?
- deleted 2y ago[deleted]
- ta988 2y agoYou will have to explain to me how.
- hbbio 2y ago
- amelius 2y agoWhy does it matter if you can run the LLM locally, if you're still running it on someone else's locked down computing platform?
- PeterStuer 2y agoRunning locally, your data is not sent outside of your security perimeter off to a remote data center. If you are going to argue that the OS or even below that the hardware could be compromised to still enable exfiltration, that is true, but it is a whole different ballgame from using an external SaaS no matter what the service guarantees.
- tempest_ 2y agoNvidia has had the Grace Hoppers for a while now. Is this not like that?
- ykl 2y agoThis is cheap compared to GB200, which has a street price of >$70k for just the chip alone if you can even get one. Also GB200 technically has only 192GB per GPU and access to more than that happens over NVLink/RDMA, whereas here it’s just one big flat pool of unified memory without any tiered access topology.
- rbanffy 2y agoWe finally encountered the situation where an Apple computer is cheaper than its competition ;-) All joking aside, I don't think Apples are that expensive compared to similar high-end gear. I don't think there is any other compact desktop computer with half a terabyte of RAM accessible to the GPU.
- kridsdale1 2y agoAnd yet all that cash still just goes to TSMC
- rbanffy 2y agoThey are selling the shovels for this gold rush. Also, ASML, who sells machines to make shovels.
- nightski 2y agoI mean expensive relative to who, Nvidia? Both are enjoying little to no competition in their respective niche and are using that monopoly power to extract massive margins. I have no doubt it could be much cheaper if there was actual competition in the market. Fortunately it seems like AMD is finally catching on and working towards producing a viable competitor to the M series chips.
- TheRealPomax 2y agoI think the other big thing is that the base model finally starts at a normal amount of memory for a production machine. You can't get less than 96GB. Although an extra $4000 for the 512GB model seems Tim Apple levels of ridiculous. There is absolutely no way that the different costs anywhere near that much at the fab. And the storage solution still makes no sense of course, a machine like this should start at 4TB for $0 extra, 8TB for $500 more, and 16TB for $1000 more. Not start at a useless 1TB, with the 8TB version costing an extra $2400 and 16TB a truly idiotic $4600. If Sabrent can make and sell 8TB m.2 NVMe drives for $1000, SoC storage should set you back half that, not over double that.
- jjtheblunt 2y ago> There is absolutely no way that the different costs anywhere near that much at the fab. price premium probably, but chip lithography errors (thus, yields) at the huge memory density might be partially driving up the cost for huge memory.
- TheRealPomax 2y agoIt's Apple, price premium is a given.
- MBCook 2y agoThis is also a niche product. The number they sell is going to be very tiny compared to the base model MacBook, let alone the iPhone. Apple absolutely loves to gouge for upgrades, but the chips in this have got to be expensive. I almost wonder if the absolute base model of this machine has much noticeably lower margins than a normal Apple product because that. But they expect/know that most everyone who buys one is going to spec it up.
- wtallis 2y ago> but chip lithography errors (thus, yields) at the huge memory density might be partially driving up the cost for huge memory. Apple's not having TSMC fab a massive die full of memory. They're buying a bunch of small dies of commodity memory and putting them in a package with a pair of large compute dies. How many of those small commodity memory dies they use has nothing to do with yield.
- deleted 2y ago[deleted]
- PeterStuer 2y agoIs this on chip memory? From the 800GB/s I would guess more likely a 512bit bus (8 channel) to DDR5 modules. Doing it on a quad channel would just about be possible, but really be pushing the envelope. Still a nice thing. As for practicality, which mainstream applications would benefit from this much memory paired with a nice but relative mid compute? At this price-point (14K for a full specced system), would you prefer it over e.g. a couple of NVIDIA project DIGITS (assuming that arrives on time and for around the announced the 3K price-point)?
- zitterbewegung 2y agoNVIDIA project DIGITS has 128 GB LPDDR5x coherent unified system memory at a 273 Gb/s memory bus speed.
- bangaladore 2y agoIt would be 273 GB/s (gigabytes, not gigabits). But in reality we don't know the bandwidth. Some ex employee said 500 GB/s. You're source is a reddit post in which they try to match the size to existing chips, without realizing that its very likely that NVIDIA is using custom memory here produced by Micron. Like Apple uses custom memory chips.
- PeterStuer 2y agoYes, but for the price of that single M3 ultra I could have 4 of those GB10's running in a 2x2 cluster with the full NVIDIA stack supported (which is still a big thing) So M3 preference will depend on whether a niche can significantly benefit from a monolitic lower compute high memory vs higher compute but distributed setup.
- MBCook 2y agoUnless something had changed its on package, but not the same die.
- sudoshred 2y agoAgree. Finally I can have several hundred browser tabs open simultaneously with no performance degradation.
- protocolture 2y agoWell at least 20
- Dban1 2y agoNew update just came in, make that 15
- nikisweeting 2y agoMy M1 Max regularly pushes 1000+ tabs without breaking a sweat, I feel like this particular metric is no longer useful now that background tab memory is almost always unloaded by the browser.
- resters 2y agoThe same thing could be designed with greater memory bandwidth, and so it's just a matter of time (for NVIDIA) until Apple decides to compete.
- RataNova 2y agoIt's a game changer for sure.... 512GB of unified memory really pushes the envelope, especially for running complex AI models locally. That said, the real test will be in how well the dual-chip design handles heat and power efficiency
- asdffdasy 2y agostill not ECC
- rlt 2y agoIs putting RAM on the same chip as processing economical? I would have assumed you’d want to save the best process/node for processing, and could use a less expensive processes for RAM.
- nullc 2y agoI'm not sure that unified memory is particularly relevant for that-- so e.g. on zen4/zen5 epyc there is more than enough arithmetic power that LLM inference is purely memory bandwidth limited. On dual (SP5) Epyc I believe the memory bandwidth is somewhat greater than this apple product too... and at apple's price points you can have about twice the ram too. Presumably the apple solution is more power efficient.