5 ms·
Every single AI shop on the planet is trying to figure out if there is enough compute or not to make this a reasonable AI path. If the answer is yes, that 10k i
by InTheArena 2y ago
Every single AI shop on the planet is trying to figure out if there is enough compute or not to make this a reasonable AI path. If the answer is yes, that 10k is a absolute bargain.
- internetter 2y agoNo AI shop is buying macs to use as a server. Apple should really release some server macOS distribution, maybe even rackable M-series chips. I believe they have one internally.
- jerjerjer 2y agoWhy would any business pay Apple Tax for a backend, server product?
- 827a 2y agoIs this actually true? Were people doing this with the 192gb of the M2 Ultra? I'm curious to learn how AI shops are actually doing model development if anyone has experience there. What I imagined was: Its all in the "cloud" (or, their own infra), and the local machine doesn't matter. If it did matter, the nvidia software stack is too important, especially given that a 512gb M3 Ultra config costs $10,000+.
- DrBenCarson 2y agoYou’re largely correct for training models Where this hardware shines is inference (aka developing products on top of the models themselves)
- 827a 2y agoTrue. But with Project Digits supposedly around the corner, which supposedly costs $3,000 and supports ConnectX and runs Blackwell; what's the over-under on just buying two of those at about half the price of one maxed M3 Ultra Mac Studio?
- DrBenCarson 2y agoAnd how much VRAM will Project Digits have?
- 827a 2y ago128gb each, so two would have 256gb. Its half that of a max spec Mac Studio, but also half the price and eight times faster memory speed. Realistically which open source LLMs does 512gb over 256gb of memory unlock? My understanding is that the true bleeding edge ones like R1 won't even handle 512gb well, especially with the anemic memory speed.
- seanmcdirmid 2y agoWe really should see what happens when Project Digits is finally released. Also, I would love in NVIDIA decided to get in the CPU/GPU + unified memory space. I can't imagine the M3 Ultra doing well on a model that loads into ~500G, but they should be a blast on 70b models (well, twice as fast as my M3 Max at least) or even a heavily quantized 400b model.
- DrBenCarson 2y agoI agree project digits looks to be the better all-around option for AI researchers, but I still think the Mac is better for people building products with AI Re memory speed, digits will be at 273GB/s while the Mac Studio is at 819GB/s Not to mention the Mac has 6 120GB/s thunderbolt 5 ports and can easily be used for video editing, app development, etc.
- Spooky23 2y ago> that 10k is a absolute bargain The higher end NVidia workstation boxes won’t run well on normal 20amp plugs. So you need to move them to a computer room (whoops, ripped those out already) or spend months getting dedicated circuits run to office spaces.
- magnetometer 2y agoDidn't really think about this before, but that seems to be mainly an issue in Northern / Central America and Japan. In Germany, for example, typical household plugs are 16A at 230V.
- someothherguyy 2y agoIn the US, normal circuits aren't always 20A, especially in residential buildings, where they are more commonly 15A in bedrooms and offices. https://en.wikipedia.org/wiki/NEMA_connector https://en.wikipedia.org/wiki/NEMA_connector
- hervature 2y agoTo clarify, the circuit is almost always 20A with 15A being used for lighting. However, the outlet itself is almost always 15A because you put multiple outlets on a single circuit. You are going to see very few 20A in outlets (which have a T shaped prong) in residential.
- badc0ffee 2y agoTo clarify further, "20A" circuit just means a 20A breaker and suitable wire (12 AWG or larger).
- theturtle32 2y agoWhile technically true, the NEMA 5-15R receptacles are rated for use on 20A circuits, and circuits for receptacles are almost always 20A circuits, in modern construction at least. Older builds may not be, of course. That said, if your load is going to be a continuous load drawing 80% of the rated amperage, it really should be a NEMA 5-20 plug and receptacle, the one where one of the prongs is horizontal instead of vertical. Swapping out the receptacle for one that accepts a NEMA 5-20P plug is like $5. If you are going to actually run such a load on a 20A circuit with multiple receptacles, you will want to make sure you're not plugging anything substantial into any of the other receptacles on that circuit. A couple LED lights are fine. A microwave or kettle, not so much.
- NorwegianDude 2y agoNot much to figure out. It's 2x M4 Max, so you need 100 of these to match the TOPS of even a single consumer card like the RTX 5090.
- jeffhuys 2y agoSure, but if you have models like DeepSeek - 400GB - that won't fit on a consumer card.
- NorwegianDude 2y agoTrue. But an AI shop doesn't care about that. They get more performance for the money by going for multiple Nvidia GPUs. I have 512 GB ram on my PC too with 8 memory channels, but it's not like it's usable for AI workloads. It's nice to have large amounts of RAM, but increasing the batch size during training isn't going to help when compute is the bottleneck.
- DrBenCarson 2y agoNow do VRAM
- wpm 2y agoIt's 2x M3 Max
- alberth 2y ago> It's 2x M4 Max Not exactly though. This can have 512GB unified memory, 2x M4 Max can only have 128GB total (64GB each).
- ZeroTalent 2y agoNo, because there is no CUDA. We have fast and cheap alternatives to NVIDIA, but they do not have CUDA. This is why NVIDIA has 90% margins on its hardware.
- jauntywundrkind 2y agoCUDA is simply not important for modern vLLM and many many others. DeepSeek V3 works great on SGLang. https://www.amd.com/en/developer/resources/technical-articles/amd-instinct-gpus-power-deepseek-v3-revolutionizing-ai-development-with-sglang.html https://www.amd.com/en/developer/resources/technical-article... Can you do absolutely everything? No. But most models will run or retrain fine now without CUDA. This premise keeps getting recycled from the past, even as that past has grown ever more distant.
- physicsguy 2y agoCUDA is incredibly important still. It's still an incredible amount of work to get packages working on multiple GPU paradigms, and by default everyone still starts with CUDA. The example I always give is FFT libraries - if you compare cuFFT to rocFFT. rocFFT only just released support for distributed transforms in December 2024, something you've been able to do since CUDA Toolkit v8.0, released in 2017. It's like this across the whole AMD toolkit, they're so far behind CUDA it's kind of laughable.
- ZeroTalent 2y agoCUDA is becoming more critical, not less, every day. Software developed around CUDA is vastly outpacing what other companies produce. And saving a few millions when creating new models doesn't matter; NVIDIA is pretty efficient at scale. I don't know if you've heard, but NVIDIA is about to add a monthly payment for additional CUDA features and I'm almost certain that many big companies will be happy to pay for them. > But most models will run or retrain fine now without CUDA. This is correct for some small startups, not big companies.