9 ms·
They didn't increase the memory bandwidth. You can get the same memory bandwidth, which is available on the M2 Studio. Yes, yes, of course you can get 512 gigab
by FloatArtifact 2y ago
They didn't increase the memory bandwidth. You can get the same memory bandwidth, which is available on the M2 Studio. Yes, yes, of course you can get 512 gigabytes of uRAM for 10 grand.
The the question is if a llm will run with usable performance at that scale? The point is there's diminishing returns despite having enough uRAM with the same amount of memory bandwidth even with increased processing speed of the new chip for AI.
So there must be a min-max performance ratio between memory bandwidth and the size of the memory pool in relation to the processing power.
- cxie 2y agoGuess what? I'm on a mission to completely max out all 512GB of mem...maybe by running DeepSeek on it. Pure greed!
- gustomksimus25 2y ago[dead]
- swivelmaster 2y agoYou could always just open a few Chrome tabs…
- DidYaWipe 2y ago[flagged]
- ksec 2y ago>Edit: WTF, someone downvoted "Enjoy the upvotes?" Pathetic. You should read HN posting Guidelines if you want to understand why. Although I guess mostly in this case it is someone fat thumbed downvote.
- School-Cotton 2y agoI downvote all Reddit-style memes, jokes, reference humor, catchphrases, and so on. It’s low-effort content that doesn’t fit the vibe of HN and actively makes the site worse for its intended purpose.
- ksec 2y agoIt may not be Firefox in terms of hundreds or thousands of tabs but Chrome has gotten a lot more memory efficient since around 2022.
- petepete 2y agoGive Cities Skylines 2 a try.
- zactato 2y agoIt doesn't support Macs yet
- deleted 2y ago[deleted]
- valine 2y agoProbably helps that models like deepseek are mixture of expert. Having all weights in VRAM means you don’t have to unlod/reload. Memory bandwidth usage should be limited to the 37B active parameters.
- FloatArtifact 2y ago> Probably helps that models like deepseek are mixture of expert. Having all weights in VRAM means you don’t have to unlod/reload. Memory bandwidth usage should be limited to the 37B active parameters. "Memory bandwidth usage should be limited to the 37B active parameters." Can someone do a deep dive above quote. I understand having the entire model loaded into RAM helps with response times. However, I don't quite understand the memory bandwidth to active parameters. Context window? How much the model can actively be processed despite being fully loaded into memory based on memory bandwidth?
- valine 2y agoWith a mixture of experts model you only need to read a subset of the weights from memory to compute the output of each layer. The hidden dimensions are usually smaller as well so that reduces the size of the tensors you write to memory.
- bick_nyers 2y agoJust to add onto this point, you expect different experts to be activated for every token, so not having all of the weights in fast memory can still be quite slow as you need to load/unload memory every token.
- valine 2y agoProbably better to be moving things from fast memory to faster memory than from slow disk to fast memory.
- ein0p 2y agoWhat people who did not actually work with this stuff in practice don't realize is the above statement only holds for batch size 1, sequence size 1. For processing the prompt you will need to read all the weights (which isn't a problem, because prefill is compute-bound, which, in turn is a problem on a weak machine like this Mac or an "EPYC build" someone else mentioned). Even for inference, batch size greater than 1 (more than one inference at a time) or sequence size of greater than 1 (speculative decoding), could require you to read the entire model, repeatedly. MoE is beneficial, but there's a lot of nuance here, which people usually miss.
- diggan 2y ago> The the question is if a llm will run with usable performance at that scale? This is the big question to have answered. Many people claim Apple can now reliably be used as a ML workstation, but from the numbers I've seen from benchmarks, the models may fit in memory, but the performance for tok/sec is so slow to not feel worth it, compared to running it on NVIDIA hardware. Although it be expensive as hell to get 512GB of VRAM with NVIDIA today, maybe moves like this from Apple could push down the prices at least a little bit.
- radlad 2y agoIt is much slower than nVidia, but for a lot of personal-use LLM scenarios, it's very workable. And it doesn't need to be anywhere near as fast considering it's really the only viable (affordable) option for private, local inference, besides building a server like this, which is no faster: https://news.ycombinator.com/item?id=42897205 https://news.ycombinator.com/item?id=42897205
- bastardoperator 2y agoIt's fast enough for me to cancel monthly AI services on a mac mini m4 max.
- staticman2 2y agoSmaller, dumber models are faster than bigger, slower ones. What model do you find fast enough and smart enough?
- TheRealPomax 2y agoYeah they did? The M4 has a max memory bandwidth of 546GBps, the M3 Ultra bumps that up to a max of 819GBps. (and the 512GB version is $4,000 more rather than $10,000 - that's still worth mocking, but it's nowhere near as much)
- okanesen 2y agoNot that dramatic of an increase actually - the M2 Max already had 400GB/s and M2 Ultra 800GB/s memory bandwidth, so the M3 Ultra's 819GB/s is just a modest bump. Though the M4's additional 146GB/s is indeed a more noticeable improvement.
- choilive 2y agoAlso should note that 800/819GB/s of memory bandwidth is actually VERY usable for LLMs. Consider that a 4090 is just a hair above 1000GB/s
- hereonout2 2y agoDoes it work like that though at this larger scale? 512GB of VRAM would be across multiple NVIDIA cards, so the bandwidth and access is parallelized. But here it looks more of a bottleneck from my (admittedly naive) understanding.
- choilive 2y agoFor inference the bandwidth is generally not parallelized because the weights need to go through the model layer by layer. The most common model splitting method is done by assigning each GPU a subset of the LLM layers and it doesn't take much bandwidth to send model weights via PCIE to the next GPU.
- manmal 2y agoMy understanding is that the GPU must still load its assigned layer from VRAM into registers and L2 cache for every token, because those aren’t large enough to hold a significant portion. So naively, for a 24GB layer, you‘d need to move up to 24GB for every token.
- lhl 2y agoSince no one specifically answered your question yet, yes, you should be able to get usable performance. A Q4_K_M GGUF of DeepSeek-R1 is 404GB. This is a 671B MoE that "only" has 37B activations per pass. You'd probably expect in the ballpark of 20-30 tok/s (depends on how much actually MBW can be utilized) for text generation. From my napkin math, the M3 Ultra TFLOPs is still relatively low (around 43 FP16 TFLOPs?), but it should be more than enough to handle bs=1 token generation (should be way <10 FLOPs/byte for inference). Now as far is its prefill/prompt processing speed... well, that's another matter.
- drited 2y agoI would be curious about context window size that would be expected when generating ballpark 20 to 20 tokens per second using Deepseek-R1 Q4 on this hardware?
- lynguist 2y agoI actually think it’s not a coincidence and they specifically built this M3 Ultra for DeepSeek R1 4-bit. They also highlight in their press release that they tested it with 600B class LLMs (DeepSeek R1 without referring to it by name). And they specifically did not stop at 256 GB RAM to make this happen. Maybe I’m reading too much into it.
- forrestthewoods 2y agoI don’t think you understand hardware timelines if you think this product had literally anything to do with anything DeepSeek.
- bustling-noose 2y agoMy thoughts too. This product was in the pipeline maybe 2-3 years ago. Maybe with LLMs getting popular a year ago they tried to fit more memory but it’s almost impossible to do that that close to a launch. Especially when memory is fused not just a module you can swap.
- bob1029 2y ago> The question is if a llm will run with usable performance at that scale? For the self-attention mechanism, memory bandwidth requirements scale ~quadratically with the sequence length.
- kridsdale1 2y agoSomeone has got to be working on a better method than that. Hundreds of billions are at stake.
- deleted 2y ago[deleted]
- deepGem 2y agoAny idea what the sRAM to uRAM ratio is on these new GPUs ? If they have meaningfully higher sRAM than the Hopper GPUs, it could lead to meaningful speedups in large model training. If they didn't increase the memory bandwidth, then 512GB will enable longer context lengths and that's about it right? No speedups For any speedups You may need some new variant of FlashAttention3 or something along similar lines to be purpose built for Apple GPUs.