6 ms·
Nvidia is in an excellent position - they have CUDA as you point out, and they are moving that into server room dominance in this application space Google has
by imwithstoopid 4y ago
Nvidia is in an excellent position - they have CUDA as you point out, and they are moving that into server room dominance in this application space
Google has TPUs but have these even made a tiny dent in Nvidia's position?
I assume anything Apple is cooking is using Nvidia in the server room already
Intel seems completely absent from this market
AMD seems content to limit its ambitions to punching Intel
its Nvidia's game to lose at this point...I wonder when they start moving in the other direction and realize they have the power to introduce their own client platform (I secretly wish they would try to mainstream a linux laptop running on Nvidia ARM but obviously this is just a fantasy)
if anything, I think Huang may not be ambitious enough!
- jitl 4y agoTegra & later Shield were attempts to get closer to full end user platform. The Nintendo Switch is their most successful such device — with a 2-year old Tegra SKU at launch. But going full force into consumer tech is a distraction for them right now. Even the enthusiast graphics market, which should be high margin, is losing their interest. They make much more selling to the big enterprise customer CEO Jensen mentions in the open paragraph.
- echelon 4y agoGamers are going to be so pissed. They subsided the advance in GPU compute and will now be ignored for the much more lucrative enterprise AI customers. Nvidia is making the right call, of course.
- smoldesu 4y agoGamers are in heaven right now. Used 30-series cards are cheap as dirt, keeping the pressure on Intel/AMD/Apple to price their GPUs competitively. The 40-series cards are a hedged bet against anything their competitors can develop - manufactured at great cost on TSMC's 4nm node and priced out-of-reach for most users. Still, it's clear that Nvidia isn't holding out their best stuff, just charging exorbitant amounts for it.
- layoric 4y agoWhere are these cheap as dirt 30 series? A 10gb 3080 is still over $500 usd used ($750 aud) when I’ve looked. When did secondhand GPUs that still cost the same as a brand new PS5 start to be considered cheap?
- my123 4y agoA PS5 is _significantly_ slower than a 3080, it's more RTX 2070 tier.
- layoric 4y agoYes, but it is a complete gaming system. My point is $500 is still a lot of money for a used GPU that is nearly 3 years old. Jut retailed brand new for $699. So yes, prices were crazy over that period for a variety of reasons, but that shouldn’t shift the gaming value proposition so dramatically.
- smoldesu 4y ago> My point is $500 is still a lot of money for a used GPU that is nearly 3 years old. Yeah. It's good hardware. You can get cheaper cards (even cost-competitive options) on PC but Nvidia won't sell them to you. Especially not now that they're got 10 billion dollars on their TSMC tab.
- hkng5994 4y agoSure but the 3080 is a single component while the PS5 includes everything needed to run the game. The way GPU prices have inflated over the past couple generations has been absurd.
- deleted 4y ago[deleted]
- dleslie 4y agoIt's OK, gaming is also having its AI moment. I fully expect future rendering techniques to lean heavily on AI for the final scene. NeRF, diffusion models, et cetera are the thin end of the wedge.
- anonylizard 4y agoHave they ever considered that the subsidy goes the other way? The margins on an A100 card is probably 100% higher than a RTX4090. Gaming industry is also like THE first industry to be revolutionized by AI. Current stuff like DLSS and AI-accelerated path tracing are mere toys compared to what will come. Nvidia will not give up gaming. When every gamer has a Nvidia card, every potential AI developer to spring up from those gamers, will use Nvidia by default. It also helps gaming GPUs are still lucrative.
- fennecfoxy 4y agoNah, volume sales & the ability to bin. A100 would be a much more expensive product if they couldn't sell defective chips as consumer GPUs. Pretty sure that the R&D cost of the workstation cards means that those cards are technically sold at a loss with Nvidia knowing that consumer sales will make up for it.
- deleted 4y ago[deleted]
- echelon 4y ago> Gaming industry is also like THE first industry to be revolutionized by AI. That's a great counter point. > Nvidia will not give up gaming. When every gamer has a Nvidia card, every potential AI developer to spring up from those gamers, will use Nvidia by default. It also helps gaming GPUs are still lucrative. Another. But Nvidia will have a lot of balancing to do and some very thirsty competitors. Though if competition arises, that too is good for gamers.
- capableweb 4y ago> I assume anything Apple is cooking is using Nvidia in the server room already I wouldn't be so quick at assuming this. Apple already ship ML-capable chips in consumer products, and they've designed and built revolutionary CPUs in modern time. I'm of course not sure about it, but I have a feeling they are gonna introduce something that kicks up the notch on the ML side sooner or later, the foundation for doing something like that is already in place.
- imwithstoopid 4y agoApple has no present experience in building big servers (they had experience at one point, but all those people surely moved on) Mac Minis don't count Sure, they are super rich and could just buy their way into the space...but so far they are really far behind in all things AI with Siri being a punchline at this point if anything, Apple proves that money alone isn't enough
- newsclues 4y agoI assumed their server experience is still working in the iCloud division.
- capableweb 4y agoI'm no Apple fan-boy at all (closer to the opposite) so it pains me a bit to say, but they have a proven track-record of having zero experience in something, then releasing something really good in that industry. The iPhone was their first phone, and it really kicked in the smartphone race into high gear. Same for the Apple Silicon processor. And those are just two relatively recent examples.
- microtonal 4y agoThe iPhone had a lot of prehistory in Apple, from Newton to iPod. Apple Silicon alo has a long history, starting with the humble beginnings as the Apple A4 in 2010, which relied on Samsung's Hummingbird for the CPU and PowerVR for the GPU (plus they acquired PA Semi in 2008). So both are not very good examples, because they build up experience over long periods.
- losteric 4y ago> I assume anything Apple is cooking is using Nvidia in the server room already For training, sure. For inference, Apple has been in a solid competitive position since M1. LLaMa, Stable Diffusion, etc, can all run on consumer devices that my tech-illiterate parents might own.
- smoldesu 4y agoLLaMa and Stable Diffusion will run on almost any device with 4gb of free memory.
- potatolicious 4y ago> "Google has TPUs but have these even made a tiny dent in Nvidia's position?" This seems unknowable without Google's internal data. The salient question is: "how many Nvidia GPUs would Google have bought if they didn't have TPUs?" The answer is probably "a lot", but realistically we don't know how many TPUs are deployed internally and how many Nvidia GPUs it displaced.
- paulmd 4y agoTesla and the Dojo architecture is another interesting one - that's another Jim Keller project and frankly Dojo may be a little underappreciated given how everything Keller touches tends to turn into gold. https://www.nextplatform.com/2022/08/23/inside-teslas-innovative-and-homegrown-dojo-ai-supercomputer/ https://www.nextplatform.com/2022/08/23/inside-teslas-innova... Much like Google, I think Tesla realized this is a capability they need, and at the scales they expect to operate, it's cheaper than buying a whole bunch of NVIDIA product.
- abudabi123 4y agoIf I recall correct Tesla went to the Dojo because the founder's vision went way beyond where NVIDIA was at their going rate.
- rolenthedeep 4y ago> AMD seems content to limit its ambitions to punching Intel What's the deal with that anyway? A lot of people want a real alternative to Nvidia, and AMD just... Doesn't care? I guess we'll have to wait for intel to release something like CUDA and then AMD will finally do something about the GPGPU demand.
- roenxi 4y agoI was wondering the same thing and thinking about it. When AMD bought ATI they viewed the GPU as a potential differentiator on CPUs. They've invested a lot of effort into CPU-GPU fusion with their APU products. That has the potential to start paying off in a big way sometime - especially if they figure our how to fuse high end GPU and CPU and just offer a GPGPU chip to everyone. I can see why AMD might put their bets here. But the trade off was that Nvidia put a lot of effort in doing linear algebra quickly and easily on their GPUs and AMD doesn't have a response to that. Especially since they probably strategised on BLAS on an APU. But it turns out there were a lot of benefits to fast BLAS and Nvidia is making all the money from that. In short, Nvidia solved a simpler problem that turned out to be really valuable, it would take AMD a long time to organise to do the same thing and it may be a misfit in their strategy. Hence ROCm sucks and I'm not part of the machine learning revolution. :(
- paulmd 4y agoAMD's graphics R&D is driven by consoles - literally. Microsoft and Sony pay huge sums in early-stage R&D and they get to set the direction of the R&D as a result. RDNA was run explicitly from the start as a semi-custom project (much to the chagrin of Raja Koduri, as this was not his fief). So was RDNA2. https://www.pcgamesn.com/amd-sony-ps5-navi-affected-vega https://www.pcgamesn.com/amd-sony-ps5-navi-affected-vega https://www.pcgamesn.com/amd/rdna-2-sony-ps5-gpu-pc https://www.pcgamesn.com/amd/rdna-2-sony-ps5-gpu-pc As such, if the console market doesn't want it, it doesn't get built. AMD is not willing to put its own money into graphics research. AMD does not really have the marketshare to get the PC market to adopt AMD-backed features that use accelerators that aren't present in the consoles. If AMD takes 20% of the market in a given year, and the PC market turns over every 6 years, this hardware support would be present in 0-3% of the PC market and 0% of the console market. So even if RDNA3 had a magic "DLSS-level" improvement that relied on some unique new accelerator they'd added in RDNA3, it'd be an uphill fight to get it adopted. Nor is AMD going to spend the money to just implement a bunch of software features anyway - they only even invested in FSR2 after it became a competitive disadvantage for them not to have something. They won't even go the 16-series vs 20-series route of having consoles be a basic architecture (with size-reduced implementations of features) and then a full-size/higher-performance implementations on PC dGPUs with more full-fledged accelerators bolted on/etc. For example they could have done this with the ML accelerators on RDNA3 - they have a slower (microcoded?) ML instruction in the basic RDNA3, and they could have thrown a more full-fledged implementation into dGPU implementations where there's more space to spare. But it's just not worth spending on any of that for them - it's a lot of R&D for a fairly narrow slice of the market that would be impacted. https://www.anandtech.com/show/13973/nvidia-gtx-1660-ti-review-feat-evga-xc-gaming/2 https://www.anandtech.com/show/13973/nvidia-gtx-1660-ti-revi... So yeah I mean she's just not that into you. Consoles set the direction of their graphics R&D. They'll tap a few other lucrative markets like HPC but they're not going to make big spends that don't have obvious ROI involved, and AMD doesn't really have the PC-gaming marketshare to care about dGPUs as an independent market worthy of R&D. People ask "why does Intel need anything except iGPUs" and for AMD the question is "why do they need anything except consoles". The rest is interesting in a "someday" sense and potentially strategically important, but day-to-day it's pretty obvious which verticals are bringing in the bacon. And for NVIDIA that's both dGPUs and datacenter - they still make a lot of money from consumer gaming, and it gives a foothold for development to progress from curiosity to research project to business deployment. AI accelerators and CUDA being on consumer hardware has been a huge boon to R&D (contrast ROCm/HIP being essentially unusable outside enterprise hardware) and the commercial market has found uses for RT cores as well. Because NVIDIA had the realization, a lot of years ago, that they are in fact a software company, that writes the software that sells the hardware. People mocked Jensen for that for a lot of years, but he was completely right and that's why he's succeeded while AMD has spun their wheels on GPGPU for 15 years now. And the problem for AMD is, consoles won't pay for a 5% more expensive chip based on blue-sky prospects of something maybe being useful in 3+ years. Or at least not unless it gets an internal backer, like DirectStorage/RDMA obviously has been adopted despite an extremely slow burn on actual usage. Optical Flow Accelerator is probably the most recent iteration of this - GCN actually had this capability as "Fluid Motion" accelerator but consoles wanted it taken back out, because it was wasted space. Now it's the underpinning of DLSS3 and likely future work in DLSS4 - the principles of "variable temporal+spatial rate shading" AMD outlines in their recent GTC presentation seem like an obvious "DLSS2 for DLSS3". I have also spoken about this idea before and I think that is where NVIDIA is going with DLSS4, but AMD has to do it without the hardware optical flow engine (except on older GCN cards ironically). https://gpuopen.com/gdc-presentations/2023/GDC-2023-Temporal-Upscaling.pdf https://gpuopen.com/gdc-presentations/2023/GDC-2023-Temporal...
- danpalmer 4y ago> I assume anything Apple is cooking is using Nvidia in the server room already I don't think Apple's server side is big or interesting. Far more interesting is the client side, because it's 1bn devices, and they all run custom Apple silicon for this. Similarly Google has Tensor chips in end user devices. Nvidia doesn't have a story for edge devices like that, and that could be the biggest issue here for them.
- paulmd 4y ago> I don't think Apple's server side is big or interesting. Far Tangent but I wish they would! Apple's e-cores would be great for servers, and they are very area/transistor efficient (even considering the node). 0.69mm2 for something with (broad strokes) Gracemont-ish performance/skylake-ish performance (but no SMT) is really good even considering the 5nm node. I think the real-world density shrinks on 5nm ended up being around 60%... so 0.69mm2 for Blizzard is like 1.1mm2 equivalent on 7nm and Avalanche is 4.1mm2, versus Zen3 at 3.1mm2 and Zen2 at 2.72mm2. 10ESF density is supposed to be similar to TSMC 5nm, dunno how true that really is in practice on actual products. But on paper that means you have Gracemont at 1.7mm vs Blizzard at 0.69mm2 and Golden Cove at 5.55mm2 vs Avalanche at 2.55mm2. Or comparing to AMD using the 1.6x conversion factor, that gives you a 7nm-area-equivalent (assuming 5nm density on 10ESF) of 2.72mm2 for Gracemont (vs Zen2 at 2.72mm2) and 8.88mm2 versus 3.1mm2 for Zen2. And that's why they're doing e-cores, and AMD is just squeezing the last little bit of space out of their existing uarch, lol. https://www.reddit.com/r/hardware/comments/qlcptr/m1_pro_10core_soc_pips_m1_max_to_head_passmarks/hj6gmb3/ https://www.reddit.com/r/hardware/comments/qlcptr/m1_pro_10c... The M1 Pro/Max dies are mostly consumed by a gigantic iGPU (in a way it's similar to the latter days of Intel quadcore era) but the cores themselves are actually quite svelte - it's actually not a case of Apple "just throwing more transistors at it", sure they are doing that in the GPU but the CPU cores themselves are very area-efficient (again, even considering the node). https://en.wikichip.org/wiki/File:kaby_lake_(dual_core)_(annotated).png https://en.wikichip.org/wiki/File:kaby_lake_(dual_core)_(ann... https://en.wikichip.org/wiki/File:kaby_lake_r_die_shot_(annotated).png https://en.wikichip.org/wiki/File:kaby_lake_r_die_shot_(anno... A Sierra Forest-style product with multiple chiplets full of nothing but e-cores would be a fantastic thing. I completely agree that Apple doesn't have any notable presence in server, but, you could make some real good products with the pieces Apple has already demonstrated. I don't have an exact source, but I recall the Asahi folks saying that based on their reverse engineering, Ultra/2-chiplets isn't the limit, the architecture is laid out to go higher on chiplets (I want to say 4 or 8) and they just aren't exploiting it right now.