6 ms·
> But inference is unique because its performance scales with high memory throughput, and you can’t assemble that by wiring together off the shelf parts in a co
by whywhywhywhy 7mo ago
> But inference is unique because its performance scales with high memory throughput, and you can’t assemble that by wiring together off the shelf parts in a consumer form factor.
Nvidia outperforms Mac significantly on diffusion inference and many other forms. It’s not as simple as the current Mac chips are entirely better for this.
- rafram 7mo agoBut where are you going to find an Nvidia GPU with 128+ GB of memory at an enthusiast-compatible price?
- ricardobayes 7mo agoThat might even be true, but how large is the TAM for such machines?
- edelans 7mo agoand let alone competing on the energy consumption!
- sippeangelo 7mo agoSome Chinese sources sell modded Nvidia GPUs with extra VRAM. They're quite affordable in comparison to even a Mac Pro.
- giwook 7mo agoAnd how much do you trust Chinese hardware?
- x______________ 7mo agoWhen there's no one left to trust, maybe you need to re-evaluate your criteria.
- sgc 7mo agoI wouldn't say that's true or even likely. It's completely possible to be in a pit of vipers where every single snake is venomous, and that is pretty much what we are seeing: With technological advances, there is a certain subset of people that will use them primarily to solidify their power and control over others. There is no utopian society right now whose government doesn't look to spy through technology, which of course is best set up at time of manufacture.
- x______________ 7mo agoAgreed. Unless you have full control over the production chain to fully produce a device, you are subject to the whims and desires of those who preside over such technological feats that we take for granted in our daily lives. To the original point, it's safe to say that highlighting a nationality with regards to trust is baseless and without merit, as would be for any other topic (men/women from x are y, z food is better here, etc..). Real life is much more complicated and nuanced past nationalities. Some might call it FUD (fear, uncertainty and doubt) but there's always a deeper rationale at the individual level as well.
- sgc 7mo agoRather than people being wary of Chinese in general, it's more that there is a high degree of government control exercised in China and they are known to be very strategic with long-term planning in regards to technology control both for spying and actual remote control of devices. We are all just looking for the least bad option. It's not like devices from other countries are immune, but they are often less organized so there is a better chance of avoiding the Chinese level of planned access. It does seem like pretty low risk in this specific case so I agree OP's comment was bit over the top, but I would have no way to make anything resembling even an educated guess as to how far their programs go.
- giwook 7mo agoYes, this is really what I was referring to. And the fact that the original comment I was replying to mentioned "modded Chinese hardware" from some unspecified, unvetted 3rd party which doesn't exactly fill me with confidence.
- embedding-shape 7mo agoGive that most of mine, and probably yours, and probably most of the world's computers are in fact made in China one way or another, some higher percentage than others, I'm guessing most of us trust our hardware enough to continue using it.
- giwook 7mo agoTrue. I was specifically referring to "modded Chinese hardware" from some unknown, unvetted third party versus say through a well-known brand that hopefully has its own rigorous QA and security processes in place.
- whywhywhywhy 7mo agoThe Mac is also chinese hardware
- platevoltage 7mo agoIt would be hilarious if you are using a Lenovo device right now.
- giwook 7mo agoI mean it's pretty funny that probably 90% of the things in our homes are made in China.
- krsw 7mo agoAt this point I trust them more than US or Israeli tech
- estimator7292 6mo agoWhich of your devices weren't made in China?
- nextaccountic 7mo agoAny links to them? Never heard of this..
- giobox 7mo agoIt’s been going on for a while. Search YouTube or the web for 48gb 4090 (this is one of the most popular modded Nvidia cards), Nvidia of course never officially made a 4090 with this much memory. There are some on sale via eBay right now. The memory controllers on some Nvidia gpus support well beyond the 16-24gb they shipped with as standard, and enterprising folks in China desolder the original memory chips and fit higher capacity ones.
- elorant 7mo agoGo at ebay and search for RTX 4090 48GBs. There's plenty of them with prices around $3.5k
- noboostforyou 7mo agoI've seen a guy who sells modded 2080 Ti with 22gb for $500 https://www.tomshardware.com/pc-components/gpus/chinese-workshops-recondition-nvidias-old-flagship-gaming-gpu-for-ai-rtx-2080-ti-upgraded-to-22gb-for-dollar499 https://www.tomshardware.com/pc-components/gpus/chinese-work... There's also unreleased Nvidia engineering samples of cards with doubled VRAM like this - https://www.reddit.com/r/nvidia/comments/1rczghu/update_unreleased_20gb_rtx_3080_ti_shunt_and/ https://www.reddit.com/r/nvidia/comments/1rczghu/update_unre...
- embedding-shape 7mo agoWhere are you gonna find Apple hardware with 128GB of memory at enthusiast-compatible price? The cheapest Apple desktop with 128GB of memory shows up as costing $3499 for me, which isn't very "enthusiast-compatible", it's about 3x the minimum salary in my country!
- joe_mamba 7mo ago> it's about 3x the minimum salary in my country! Enthusiast compute hardware doesn't cater to the people on the minimum salary in any country, let alone developing nations. When Ferrari makes a car they don't ask themselves if people on minimum salary will be able to afford them. In in the bottom two poorest EU member states and Apple and Microsoft Xbox don't even bother to have a direct to customer store presence here, you buy them from third party retailers. Why? Probably because their metrics show people here are too poor to afford their products en-masse to be worth operating a dedicated sales entity. Even though plenty of people do own top of the line Macbooks here, it's just the wealthy enthusiast niche, but it's still a niche for the volumes they (wish to)operate at. Why do you think Apple launched the Mac Neo?
- embedding-shape 7mo agoRight, I think maybe we're then talking about "upper class enthusiasts" or something in reality then? I understood that to juts be about the person, not what economic class they were in, maybe I misunderstood.
- deleted 7mo ago[deleted]
- Heliosmaster 7mo agoYes, it's a different definition. Enthusiast in this contest more or less means you are excited enough about something to get a level above what normal people should get and just below professional pricing. An enthusiast camera body can be 2000 euros. I would say an enthusiast computer is 2-4k. It really depends what you meant with minimum salary (yearly?) because paying 3 months of salary for a computer like that isn't far fetched. You're not using this to generate recipes for cookies. An enthusiast level car is expensive as well.
- angoragoats 7mo agoYou can still buy used 3090 cards on ebay. 5 of them will give you 120GB of memory and will blow away any mac in terms of performance on LLM workloads. They have gone up in price lately and are now about $1100 each, but at one point they were $700-800 each.
- rybosworld 7mo agoI don't see how 5x 3090's is a better option than an M3 Ultra Mac studio. The mac will just work for models as large as 100B, can go higher with quantized models. And power draw will be 1/5th as much as the 3090 setup. You can certainly daisy chain several 3090's together but it doesn't work seamlessly.
- angoragoats 7mo ago> The mac will just work for models as large as 100B, can go higher with quantized models. And power draw will be 1/5th as much as the 3090 setup. This setup will work for 100B models as well. And yes, the Mac will draw less power, but the Nvidia machine will be many times faster. So depending on your specific Mac and your specific Nvidia setup, the performance per watt will be in the same ballpark. And higher absolute performance is certainly a nice perk. > You can certainly daisy chain several 3090's together but it doesn't work seamlessly. Citation needed; there's no "daisy chaining" in the setup I describe, and low level libraries like pytorch as well as higher level tools like Ollama all seamlessly support multiple GPUs.
- lowbloodsugar 7mo agoHow much does it cost to have an electrician wire up 240v circuit just to power the thing?
- angoragoats 7mo agoThe machine I’m describing works just fine on a dedicated 15A 120V circuit.
- dabockster 7mo agoYou don’t need it if you use llamacpp on Windows, or if you compile it on Linux with CUDA 13 and the correct kernel HMM support, and you’re only using MoE models (which, tbh, you should be doing anyways).
- 0x457 7mo agoWhat MoE has to do with it? Aside from Flash-MoE that supports exactly one model and only on macOs - you still need to load entire model into memory. You also don't know what experts going to be activated, so it's not like you can predict which needs to be loaded.
- zozbot234 7mo agoWith proper mmap support you don't really need the entire model in memory. It can be streamed from a fast SSD, and this is more useful for MoE models where not all expert-layers are uniformly used. Of course the more data you stream from SSD, the slower this is; caching stuff in RAM is still relevant to good performance.
- gloxkiqcza 6mo agoYou can do this on a Mac as well tho, right? So that 128 GB unified memory becomes cache for very fast 1+ TB Apple SSD.
- zozbot234 6mo agoI think the advantage of Flash-MoE compared to plain mmap is mostly the coalesced representation where a single expert-layer is represented by a single extent of sequential data. That could be introduced to existing binary formats like GGUF or HF - there is already a provision for differently structured representations, and that would easily fit.
- 0x457 6mo agoOkay, yes, you don’t need the entire MoE model in memory for it to function. But you still need the working set of frequently used experts to actually fit in RAM, or at least stay cached. Expert routing happens per token, per layer. If those weights aren’t resident, you’re effectively pulling them from disk on the critical path of generation — over and over again. That’s not “just slower,” that’s order of magnitude slower. You’ll end up with constant page faults and page cache churn. And if swap is on the same device as the model, you’re now competing for bandwidth on top of that. IMO the main benefit of mmap is ability to reclaim cold pages during high memory-pressure events when model isn't active.
- colechristensen 7mo agoThe Nvidia DGX Spark is exactly this and in the same price and performance bracket.
- andreybaskov 7mo agoSadly, memory bandwidth is abysmal compared to Apple chips - 273 GB/s vs 614 GB/s on M5 Max for similar price. Even though fp4 compute is faster, it doesn't help for all the decode heavy agentic workflows.
- chpatrick 7mo agoBut they're pretty fast and can have loads of RAM, which would be prohibitively expensive with Nvidia.
- chocochunks 7mo agoA 128GB 2TB Dell Pro Max with Nvidia GB10 is about $4200, a Mac Studio with 128GB RAM and 2TB storage is $4100. So pretty comparable. I think Dell's pricing has been rocked more by the RAM shortage too.
- plagiarist 7mo agoNot quite, what is the vRAM bandwidth of each? The bandwidth is a huge contributor to LLM performance.
- embedding-shape 7mo agoAFAIK, for the unified bandwidth, it depends mostly on the CPU, for M4 Max (I think it's the default today?) it does ~550 GB/s, while GB10 does ~270 GB/s, so about a 2x difference between the two. For comparison, RTX Pro 6000 does 1.8 TB/s, pretty much the same as what a 5090 does, which is probably the fastest/best GPUs a prosumer reasonable could get.
- plagiarist 7mo agoGranted, it won't be competitive against the flagship dGPUs. Nevertheless, that ~2x is a pretty huge difference in similarly priced offerings.
- midnight_eclair 7mo ago~not unified memory tho~
- mciancia 7mo agoIt is unified memory on this one
- AdamN 7mo agoNvidia isn't selling one-off home computers afaik. But yes in terms of datacenter cloud usage Nvidia performs.
- newsclues 7mo agohttps://marketplace.nvidia.com/en-us/enterprise/personal-ai-supercomputers/dgx-spark/ https://marketplace.nvidia.com/en-us/enterprise/personal-ai-...
- jamespo 7mo agoAmusingly there's a macbook next to it in the pic, is this headless?
- Tsiklon 7mo agoIt has a HDMI port and its USB-C ports also support display out. But I believe most who buy it intend to use it headless. The machine runs Ubuntu 24.04 and has a slightly customised Gnome (green accents and an nvidia logo in GDM) as its desktop.
- _zoltan_ 7mo agoGB300 DGX Station was announced last Monday.
- eitally 7mo agoIt's going to cost far more than a diy machine with multiple lower end GPUs. Which is fine -- it's aimed at enterprise, not home labs.
- wappieslurkz 7mo agoDo NVIDIA solutions also outperform the Apple M-series in performance per Watt?
- Lalabadie 7mo agoProbably comparable, but that's only with business-grade products, it's why Apple's current silicon is so remarkable on the market at the consumer level.
- wappieslurkz 7mo agoThanks.
- whywhywhywhy 7mo agoNo, that's why Apple uses Performance Per Watt not actual performance celling as the metric. In actual workloads where you'd need this power then actual performance is what matters not PPW.
- jiwidi 7mo agotell me what pc with an nvidia gpu can you buy with same memory and performance. I never liked apple hardware, but they are now untouchable since their shift to own sillicon for home hardware.
- traceroute66 7mo ago> tell me what pc with an nvidia gpu can you buy with same memory and performance. And power consumption ! The performance per watt of Apple is unmatched.
- dabockster 7mo agoThis needs to be sold as the big ticket item for low level devs. Their chips are some of the most power efficient chips on the market right now. Hoping they release a blade server version somehow.
- bigyabai 7mo agoNvidia's recent GPUs are more power-efficient than Apple Silicon in raster, training and inference workloads. A blade server would get cancelled just like the Mac Pro for exactly the same reasons: https://9to5mac.com/2026/03/02/some-apple-ai-servers-are-reportedly-sitting-unused-on-warehouse-shelves-due-to-low-apple-intelligence-usage/ https://9to5mac.com/2026/03/02/some-apple-ai-servers-are-rep...
- traceroute66 7mo ago> Nvidia's recent GPUs are more power-efficient than Apple Silicon in raster, training and inference workloads. I think you can do better than the proverbial Apples and Oranges comparison. In terms of total system, "box on desk", Apple is likely to remain the performance per watt leader compared to random PC workstations with whatever GPUs you put inside.
- bigyabai 7mo ago