2 ms·
>> 252GB of HBM3e at 7.1TB/s, So it has only 252GB of actual ”AI” memory making it “useless”/toy for actual real world AI workloads(I.e it can’t replace someth
by theplumber 18d ago
>> 252GB of HBM3e at 7.1TB/s,
So it has only 252GB of actual ”AI” memory making it “useless”/toy for actual real world AI workloads(I.e it can’t replace something like opus 5)
- kees99 18d agoShould work just fine for MoE models where active set fits into 252GB.
- theplumber 18d agoCan’t you do that already more or less with a Mac Studio with 256 or better 512fb of ram?
- fc417fc802 18d agoWhere are you getting 7 TB/s of memory bandwidth?
- theplumber 18d agoFair point but you get that only for 252GB of ram so it’s not even “totally better” than a Mac Studio ultra with 512GB RAM. It still the old “smaller but faster than a Mac” stuff nvidia sells.Maybe in 1-2 years they will match the Mac 512 but then Apple will release an even bigger Mac
- fc417fc802 18d agoI don't know about "useless" (it seems quite useful to me) but I do feel mislead. It's unified memory in the same way that my current dGPU has unified memory. I guess nvlink-c2c probably (?) doesn't introduce a bottleneck but it's still two distinct arenas with very different performance characteristics.
- theplumber 18d agoI find it just misleading by advertising 700GB RAM as AI headline. I could plug a 32GB GPU to my 1.5TB ram server and call it “AI station with 1.5TB+ RAM” just so that you find it actually useless compared to the headline
- cmrdporcupine 18d ago7.1TB/s of HBM is not "useless" -- that's 30 times the memory bandwidth of my DGX Spark -- and nobody is expecting such a machine to run "Opus 5" on its own. For such large models even datacentre GB300 NVL72 are multiple trays linked together via NVlink etc. This machine has QSFP ports and ConnectX for linking up for larger models. It's a workstation, not a rack. It's for AI researchers. I'd love to have one on (err, under) my desk. What even is this comment?
- theplumber 18d agoMy point is this is “useless” because you can’t run large models. It’s like having a very fast and expensive SSD with a very small capacity. It’s not useless but it is for “real work”. Apple has been providing 512GB ram machines for several years so to say you need a rack to run a large model for personal usage it’s missing the point. You needed a rack of nvidia cards to match an 128GB ram Mac as well a while ago so it’s more of the same.
- cmrdporcupine 18d agoPeople who buy Macs to run local inference are not the target audience for this machine. It's an AI research workstation for people whose ultimate work goes on production GB300 NVL72 data centre racks (and for large models you link them together.) It's also about 15x the memory bandwidth of a Mac. For models that fit you'd be looking at hundreds of tokens per second on decode and prefill many times that. And has CUDA, which is (likely) what your production system will use. And runs a real server operating system. Also by the time you spec'd out a Mac with the same total memory capacity and computation you'd also be as expensive. And still not have as many cores, nor have the ConnectX RDMA networking speeds.
- LargoLasskhyfv 18d ago[dead]