8 ms·
Microsoft CTO says he wants to swap most AMD and Nvidia GPUs for homemade chips
- okokwhatever 1y agosomebody wants to buy NVDA cheaper... ;)
- synergy20 1y agoplus someone has no leverage whatsoever other than talking
- ambicapter 1y agoMicrosoft, famously resource-poor.
- fidotron 1y agoThey have so much money it is harmful to their ability to execute. Just look at the implosion of the XBox business.
- fennecbutt 1y agoGranted, if everyone had done what the highly paid executives had told them to do, xbox would never have existed. And I'm guessing that the decline is due to executive meddling. What is it that executives do again? Beyond collecting many millions of dollars a year, that is.
- fidotron 1y agoThey sit around fantasizing about buying Nintendo because that would be the top thing they could achieve in their careers.
- deleted 1y ago[deleted]
- rjbwork 1y agoGuess MSFT needs somewhere else AI adjacent to funnel money into to produce the illusion of growth and future cash flow in this bubblified environment.
- balls187 1y ago> produce the illusion of growth and future cash flow in this bubblified environment. I was ranting about this to my friends; Wallstreet is now banking on Tech firms to produce the illusion of growth and returns, rather than repackaging and selling subprime mortgages. The tech sector seems to have a never ending supply of things to spur investment and growth: cloud computing, saas, mobile, social media, IoT, crypto, Metaverse, and now AI. Some useful, some not so much. Tech firms have a lot of pressure to produce growth, it's filled with very smart people, and wields influence on public policy. The flip side is the mortage crisis, at least before it collapsed, got more Americans into home ownership (even if they weren't ready for it). I'm not sure the tech sectors meteoric rise has been as helpful (sentiment of locals in US tech hubs suggests a overall feeling of dissatisfaction with tech)
- giancarlostoro 1y agoSo similar to Apple Silicon. If this means they'll be on par with Apple Silicon I'm okay with this, I'm surprised they didn't do this sooner for their Surface devices. Oh right, for their data centers. I could see this being useful there too, brings costs down lower.
- CharlesW 1y ago> So similar to Apple Silicon. Yes, in the sense that this is at least partially inspired by Apple's vertical integration playbook, which has now been extended to their own data centers based on custom Apple Silicon¹ and a built-for-purpose, hardened edition of Darwin². ¹ https://security.apple.com/blog/private-cloud-compute/ https://security.apple.com/blog/private-cloud-compute/ ² https://en.wikipedia.org/wiki/Darwin_(operating_system) https://en.wikipedia.org/wiki/Darwin_(operating_system)
- giancarlostoro 1y agoYeah, its interesting, years ago I never thought Apple nor Microsoft would do this, but also Google has done this on their cloud as well, so it makes sense.
- georgeburdell 1y agoVertical integration only works if your internal teams can stay in the race at each level well enough to keep the stack competitive as a whole. Microsoft can’t attract the same level of talent as Apple because their pay is close to the industry median
- quadrature 1y agoNot suprising that the hyperscalers will make this decision for inference and maybe even a large chunk of training. I wonder if it will spur nvidia to work on an inference only accelerator.
- edude03 1y ago> I wonder if it will spur nvidia to work on an inference only accelerator. Arguably that's a GPU? Other than (currently) exotic ways to run LLMs like photonics or giant SRAM tiles there isn't a device that's better at inference than GPUs and they have the benefit that they can be used for training as well. You need the same amount of memory and the same ability to do math as fast as possible whether its inference or training.
- conradev 1y agoThey’re already optimizing GPU die area for LLM inference over other pursuits: the FP64 units in the latest Blackwell GPUs were greatly reduced and FP4 was added
- CharlesW 1y ago> Arguably that's a GPU? Yes, and to @quadrature's point, NVIDIA is creating GPUs explicitly focused on inference, like the Rubin CPX: https://www.tomshardware.com/pc-components/gpus/nvidias-new-cpx-gpu-aims-to-change-the-game-in-ai-inference-how-the-debut-of-cheaper-and-cooler-gddr7-memory-could-redefine-ai-inference-infrastructure https://www.tomshardware.com/pc-components/gpus/nvidias-new-... "…the company announced its approach to solving that problem with its Rubin CPX— Content Phase aXcelerator — that will sit next to Rubin GPUs and Vera CPUs to accelerate specific workloads."
- edude03 1y agoYeah, I'm probably splitting hairs here but as far as I understand (and honestly maybe I don't understand) - Rubin CPX is "just" a normal GPU with GDDR instead of HBM. In fact - I'd say we're looking at this backwards - GPUs used to be the thing that did math fast and put the result into a buffer where something else could draw it to a screen. Now a "GPU" is still a thing that does math fast, but now sometimes, you don't include the hardware to put the pixels on a screen. So maybe - CPX is "just" a GPU but with more generic naming that aligns with its use cases.
- hkt 1y agoEven just saying this applies downward pressure on pricing: NVIDIA has an enormous amount of market power (~"excess" profit) right now and there aren't enough near competitors to drive that down. The only thing that will work is their biggest _consumers_ investing, or threatening to invest, if their prices are too high. Long term, I wonder if we're exiting the "platform compute" era, for want of a better term. By that I mean compute which can run more or less any operating system, software, etc. If everyone is siloed into their own vertically integrated hardware+operating system stack, the results will be awful for free software.
- startupsfail 1y agoIn that case, it's great that Microsoft is building their silicon. Keeps NVIDIA in check, otherwise these profits would evaporate into nonsense and NVIDIA would lose the AI industry to competition from China. Which, depending if AGI/ASI is possible or not, may or may not be a great move.
- harrall 1y agoGoogle has been using its own TPU silicon for machine learning since 2015. I think they do all deep learning for Gemini on ther own silicon. But they also invented AI as we know it when they introduced transformer architecture and they’ve been more invested in machine learning than most companies for a very long time.
- chrismustcode 1y agoI thought they use GPU for learning and TPU for inference, I’m open to been corrected.
- dekhn 1y agono. for internal training most work is done on TPUs, which have been explicitly designed for high performance training.
- cendyne 1y agoI've heard its a mixture because they can't source enough in-house compute
- xnx 1y agoSome details here: https://news.ycombinator.com/item?id=42392310 https://news.ycombinator.com/item?id=42392310
- surajrmal 1y agoThe first tpu they made was inference only. Everything since has been used for training. I think that means they weren't using it for training in 2015 but rather 2017 based on Wikipedia.
- lokar 1y agoThe first TPU they *announced" was for inference
- buildbot 1y agoNot that it matters, but Microsoft has been doing AI accelerators for a bit too - project Brainwave has been around since 2018 - https://blogs.microsoft.com/ai/build-2018-project-brainwave/ https://blogs.microsoft.com/ai/build-2018-project-brainwave/
- tetrisgm 1y agoHonestly this would be great for competition. Would love to see them impish in that direction.
- okkdrjfhfn 1y ago[flagged]
- bee_rider 1y agoThe most important note is: > The software titan is rather late to the custom silicon party. While Amazon and Google have been building custom CPUs and AI accelerators for years, Microsoft only revealed its Maia AI accelerators in late 2023. They are too late for now, they realistically hardware takes a couple generations to become a serious contender and by the time Microsoft has a chance to learn from their hardware mistakes the “AI” bubble will have popped. But, there will probably be some little LLM tools that do end up having practical value; maybe there will be a happy line-crossing point for MS and they’ll have cheap in-house compute when the models actually need to be able to turn a profit.
- surajrmal 1y agoAt this point it will take a lot of investment to catch up. Google relies heavily on specialized interconnects to build massive tpu clusters. It's more than just designing a chip these days. Folks who work on interconnects are a lot more rare than engineers who can design chips.
- alephnerd 1y ago> hardware takes a couple generations to become a serious contender Not really and for the same reason Chinese players like Biren are leapfrogging - much of the workload profile in AI/ML is "embarrassingly parallel", thus reducing the need for individual ASICs to be bleeding edge performant. If you are able to negotiate competitive fabrication and energy supply deals, you can mass produce your way into providing "good enough" performance. Finally, the persona who cares about hardware performance in training isn't in the market for cloud offered services.
- kenjackson 1y agoAnd current LLM architectures affinitize differently to HW than DNNs even just a decade ago. If you have the money and technical expertise (both of which I assume MS has access to) then a late start might actually be beneficial.
- rcxdude 1y agoAs I understood it the main bottleneck is interconnects, anyhow. It's more difficult to keep the ALUs fed than it is to make them fast enough, especially once your model can't fit in one die/PCB. And that's in principle a much trickier part of the design, so I don't really know how that shakes out (is there a good enough design that you can just buy as a block?)
- cjbgkagh 1y agoI guess Microsoft’s investment into Graphcore didn’t pay off. Not sure what they’re planning but more of that isn’t going to cut it. At the time (late 2019) I was arguing for either a GPU approach or specialized architecture targeting transformers. There was a split at MS where the ‘Next Gen’ bayesian was being done in the US and the frequentist work was being shipped off to China. Chris Bishop was promoted to head of MSR Cambridge which didn’t help. Microsoft really is an institutionally stupid organization so I have no idea on which direction they actually go. My best guess is that it’s all talk.
- Den_VR 1y agoMicrosoft lacks the credibility and track record for this to be anything but talk. Hardware doesn’t simply go from zero to gigawatts of infrastructure on talk. Even Apple is better positioned for such a thing.
- migueldeicaza 1y agoThey do have such a dedicated chip, the MAIA 100 chip which is an in-house chip, and it is a chip that was designed in the era of transformers, and this is what is being discussed in the interview.
- cjbgkagh 1y agoI missed that, it’s been a few years since I’ve paid attention to MS hardware and it is very possible that my thoughts are out of date. I left MS with a rather bad taste in my mouth. I’m checking out the info on that chip and what I am seeing is a little light on details. Just TPUs and fast interconnects. What I’ve found; MIAI 200 the next version is having issues due to brain drain, and MIAI 300 is to be an entirely new architecture so the status for that is rather uncertain. I think a big reason MS invested so heavily into OpenAI was to have a marquee customer push cultural change through the org, which was a necessary decision. If that eventually yields in a useful chip I will be impressed, I hope it does.
- latchkey 1y agoIt always falls back on the software. AMD is behind, not because the hardware is bad, but because their software historically has played second fiddle to their hardware. The CUDA moat is real. So, unless they also solve that issue with their own hardware, then it will be like the TPU, which is limited to usage primarily at Google, or within very specific use cases. There are only so many super talented software engineers to go around. If you're going to become an expert in something, you're going to pick what everyone else is using first.
- amelius 1y ago> The CUDA moat is real. I don't know. The transformer architecture uses only a limited number of primitives. Once you have ported those to your new architecture, you're good to go. Also, Google has been using TPUs for a long time now, and __they__ never hit a brick wall for a lack of CUDA.
- latchkey 1y agoIt is beyond porting, it is mentality of developers. Change is expensive and I'm not just talking about $ value. > Also, Google has been using TPUs for a long time now, and __they__ never hit a brick wall for a lack of CUDA. That's exactly what I'm saying. __they__ is the keyword.
- amelius 1y agoNot sure what you mean. Google is a big company. Their TPUs have many users internally.
- latchkey 1y agoVery few developers outside of Google have ever written code for a TPU. In a similar way, far fewer have written code for AMD, compared to NVIDIA. If you're going to design a custom chip and deploy it in your data centers, you're also committing to hiring and training developers to build for it. That's a kind of moat, but with private chips. While you solve one problem (getting the compute you want), you create another: supporting and maintaining that ecosystem long term. NVIDIA was successful because they got their hardware into developers hands, which created a feedback loop, developers asked for fixes/features, NVIDIA built them, the software stack improved, and the hardware evolved alongside it. That developer flywheel is what made CUDA dominant and is extremely hard to replicate because the shortage of talented developers is real.
- outside1234 1y agoFor GPUs at least this is pretty obvious. For CPUs it is less clear to me that they can do it more efficiently.
- sgerenser 1y agoWhen Microsoft talks about “making their own CPUs,” they just mean putting together a large number of off-the-shelf Arm Neoverse cores into their own SoC, not designing a fully custom CPU. This is the same thing that Google and Amazon are doing as well.
- qwertytyyuu 1y agobetter late than never to get into to game... right? right....?
- johncolanduoni 1y agoJust like mobile!
- alephnerd 1y agoI've mentioned this before on HN [0][1]. The name of the game has been custom SoCs and ASICs for a couple years now, because inference and model training is an "embarrassingly parallel" problem, and models that are optimized for older hardware can provide similar gains to models that are run on unoptimized but more performant hardware. Same reason H100s remain a mainstay in the industry today, as their performance profile is well understood now. [0] - https://news.ycombinator.com/item?id=45275413 https://news.ycombinator.com/item?id=45275413 [1] - https://news.ycombinator.com/item?id=43383418 https://news.ycombinator.com/item?id=43383418
- philipwhiuk 1y ago> The name of the game has been custom SoCs and ASICs for a couple years now, because inference and model training is an "embarrassingly parallel" problem, and models that are optimized for older hardware can provide similar gains to models that are run on unoptimized but more performant hardware. Is anyone else getting crypto flashbacks?
- kcb 1y agoThe difference is crypto wasn't memory and throughput dependent. That's why a small asic on a USB stick could outperform a GPU.
- alephnerd 1y agoOne thing to point out - "ASIC" is more of a business term than a technical term. The teams that work on custom ASIC design at (eg.) Broadcom for Microsoft are basically designing custom GPUs for MS, but these will only meet the requirements that Microsoft lays out, and Microsoft would have full insight and visibility into the entire architecture.
- deleted 1y ago[deleted]
- croisillon 1y agohomemade chips is probably a lot of fun but buying regular Lay's is so much easier
- appleaday1 1y agoThe current M$ sure is doing a great job at making people move to alternatives.
- bongodongobob 1y agoPeople, sure, but that's not their target demographic. It's businesses and they aren't moving away from MS anytime soon.
- floxy 1y agoOn a slightly different tangent, is anyone working on analog machine learning ASICs? Sub-threshold CMOS or something? I mean even at the research level? Using a handful of transistor for an analog multiplier. And get all of the crazy fascinating translinear stuff of Barrie Gilbert fame. https://www.electronicdesign.com/technologies/analog/article/21807652/whats-all-this-subthreshold-stuff-anyhow https://www.electronicdesign.com/technologies/analog/article... https://www.analog.com/en/resources/analog-dialogue/articles/considering-multipliers-part-1.html https://www.analog.com/en/resources/analog-dialogue/articles... http://madvlsi.olin.edu/bminch/talks/090402_atact.pdf http://madvlsi.olin.edu/bminch/talks/090402_atact.pdf
- SmoothBrain123 1y ago[dead]
- nickpsecurity 1y agoA bunch of people. Just type these terms into DuckDuckGo: analog neural network hardware physical neural network hardware Put "this paper" after each one to get academic research. Try it with and without that phrase. Also, add "survey" to the next iteration. The papers that pop up will have the internal jargon the researchers use to describe their work. You can further search with it. The "this paper," "survey," and internal jargon in various combinations are how I find most CompSci things I share.
- xadhominemx 1y agoFor large models, the bottlenecks are memory bandwidth, network, and power consumption by the DAC/ADC arrays It’s never come even close to penciling out in practice. For small models there are people working on this implemented in flash memory eg Mythic.
- improgrammer007 1y agoFor those who don't know Msft is working on https://azure.microsoft.com/en-us/blog/azure-maia-for-the-era-of-ai-from-silicon-to-software-to-systems/ https://azure.microsoft.com/en-us/blog/azure-maia-for-the-er...
- hulitu 1y ago> For those who don't know Msft is working on We do know: ads, spyware and rounding corners of UI elements. If their processor work like their software, i really feel pity for people who use it.
- snowwrestler 1y agoMade where? Isn’t foundry capacity the limiting factor on chips for AI right now?
- Jyaif 1y agoBy cutting the middle man out, MS could pay TSMC more than nvidia per wafer and still save money.
- fishmicrowaver 1y agoThis is the whole game right here.
- nsteel 1y agoThey can even pay Broadcom to be a lower-level middle man instead. Despite the BRCM tax, it'll still be way cheaper than going to Nvidia.
- migueldeicaza 1y agoTSMC manufactures the MAIA 100: https://azure.microsoft.com/en-us/blog/azure-maia-for-the-era-of-ai-from-silicon-to-software-to-systems/ https://azure.microsoft.com/en-us/blog/azure-maia-for-the-er...
- blibble 1y agowell yeah, I can't imagine sending all your shareholders money to nvidia to produce slop no-one is willing to pay for is going down too well
- sleepybrett 1y agoMicrosoft just can't stop following apple's lead.
- pjmlp 1y agoDoesn't come as a surprise, I imagine they would also build on top of Direct Compute, or something else they can think of.
- alecco 1y agoFor many years, every few months Microsoft and Meta say they are going to do AI hardware. But nothing tangible is delivered.
- jtfrench 1y ago"Microsoft Silicon", coming up. Is it practical for them to buy an existing chip maker? Or would they just go home-grown? - As of today, Nvidia's market cap is a whopping 4.51 trillion USD compared to Microsoft's 3.85 trillion USD, so that might not work. - AMD's market cap is 266.49 billion USD, which is more in reach.
- floxy 1y agoIntel is almost a bargain at $175 billion
- j_walter 1y agoThey were more of a bargain 3 months ago at $80B.
- hulitu 1y ago> "Microsoft Silicon", coming up. Will they equip the new Microsoft Vacuum Cleaner with it ? /s
- arisAlexis 1y agoThey could buy with peanut money Cerebras
- foobarbecue 1y agoWeird use of "homemade"! I guess they mean "in-house"?
- ravenstine 1y agoJust like how mama used to make 'em!
- brnt 1y agoMoms secret ingredient to her AI was Nvidia!
- smallmancontrov 1y agoIt's a delightful coincidence of history that the "What's Jensen been cooking" pandemic gag happened on the generation that would wake up the AIs. https://www.youtube.com/watch?v=So7TNRhIYJ8 https://www.youtube.com/watch?v=So7TNRhIYJ8
- TacticalCoder 1y ago[dead]
- john01dav 1y agoI don't like this trend where all the big tech companies are bringing hardware in-house, because it makes it unavailable to everyone else. I'd rather not have it be the case that everyone who isn't big tech either pays the Nvidia tax or deals with comparatively worse hardware, while each big tech has their own. If these big tech companies also sold their in-house chips, then that would address this problem. I like what Ampere is doing in this space.
- Etheryte 1y agoAnother way to look at it would be that if all the big players stop buying up all the GPUs, prices will come back down for regular consumers, making it available for everyone else.
- jlarocco 1y agoBut there's a real risk that in the long term all of the "serious" AI hardware research will get done inside a few big companies, essentially shutting out smaller players with a hardware moat. Unless they start selling the hardware, but in the current AI market nobody would do that because it's their special sauce. On the other hand, maybe it's no different than any other hardware, and other makers will catch up eventually.
- dang 1y agoUrl changed from https://www.theregister.com/2025/10/02/microsoft_maia_dc/ https://www.theregister.com/2025/10/02/microsoft_maia_dc/, which points to this. Submitters: "Please submit the original source. If a post reports on something found on another site, submit the latter." - https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- eptcyka 1y agoI do too.
- yalogin 1y agoThe big data center owners will want to build their own hardware, that is a no brainer for them.
- cmxch 1y agoThen let’s hope it translates to more end consumer GPU availability for compute versus proprietary silicon.