19 ms·
TPUs vs. GPUs and why Google is positioned to win AI race in the long term
- lvl155 10mo agoRight because people would love to get locked into another even more expensive platform.
- svantana 10mo agoThat's mentioned in the article, but is the lock-in really that big? In some cases, it's as easy as changing the backend of your high-level ML library.
- LogicFailsMe 10mo agoThat's what it is on paper. But in practice you trade one set of hardware idiosyncrasies for another and unless you have the right people to deal with that, it's a hassle.
- lvl155 10mo agoOn top, when you get locked into Google Cloud, you’re effectively at the mercy of their engineers to optimize and troubleshoot. Do you think Google will help their potential competitors before they help themselves? Highly unlikely considering their actions in the past decade plus.
- LogicFailsMe 10mo agoGiven my Fitbit's inability to play nice with my pixel phone, I have zero faith in Google engineers. What else would one expect when their core value is hiring generalists over specialists* and their lousy retention record? *Pay no attention to the specialists they acquihire and pay top dollar... And even they don't stick around.
- Irishsteve 10mo agoI thin k you can only run on google cloud not aws bare metal azure etc
- tempest_ 10mo agoThat is like how every ORM promises you can just swap out the storage layer. In practice it doesnt quite work out that way.
- sbarre 10mo agoA question I don't see addressed in all these articles: what prevents Nvidia from doing the same thing and iterating on their more general-purpose GPU towards a more focused TPU-like chip as well, if that turns out to be what the market really wants.
- blibble 10mo agothe entire organisation has been built over the last 25 years to produce GPUs turning a giant lumbering ship around is not easy
- sbarre 10mo agoFor sure, I did not mean to imply they could do it quickly or easily, but I have to assume that internally at Nvidia there's already work happening to figure out "can we make chips that are better for AI and cheaper/easier to make than GPUs?"
- coredog64 10mo agoIsn't that a bit like Kodak knowing that digital cameras were a thing but not wanting to jeopardize their film business?
- sofixa 10mo ago> what prevents Nvidia from doing the same thing and iterating on their more general-purpose GPU towards a more focused TPU-like chip as well, if that turns out to be what the market really wants. Nothing prevents them per se, but it would risk cannibalising their highly profitable (IIRC 50% margin) higher end cards.
- numbers_guy 10mo agoNothing in principle. But Huang probably doesn't believe in hyper specializing their chips at this stage because it's unlikely that the compute demands of 2035 are something we can predict today. For a counterpoint, Jim Keller took Tenstorrent in the opposite direction. Their chips are also very efficient, but even more general purpose than NVIDIA chips.
- clickety_clack 10mo agoAny chance of a bit of support for jax-metal, or incorporating apple silicon support into Jax?
- deleted 10mo ago[deleted]
- dana321 10mo agoThat and the fact they can self-fund the whole AI venture and don't require outside investment.
- jsheard 10mo agoThat and they were harvesting data way before it was cool, and now that it is cool, they're in a privileged position since almost no-one can afford to block GoogleBot. They do voluntarily offer a way to signal that the data GoogleBot sees is not to be used for training, for now, and assuming you take them at their word, but AFAIK there is no way to stop them doing RAG on your content without destroying your SEO in the process.
- deleted 10mo ago[deleted]
- lazyfanatic42 10mo agoWow, they really got folks by the short hairs if that is true...
- boredatoms 10mo agoDo people still get organic search traffic from google?
- smj-edison 10mo agoBut they also collect the data without causing denial of service, and respect robots.txt, which is more than you can say of most LLM scrapers...
- mrbungie 10mo agoThe most fun fact about all the developments post-ChatGPT is that people apparently forgot that Google was doing actual AI before AI meant (only) ML and GenAI/LLMs, and they were top players at it. Arguably main OpenAI raison d'être was to be a counterweight to that pre-2023 Google AI dominance. But I'd also argue that OpenAI lost its way.
- bhouston 10mo agoIn my 20+ years of following NVIDIA, I have learned to never bet against them long-term. I actually do not know exactly why they continually win, but they do. The main issue they have a 3-4 year gap between wanting a new design pivot and realizing it (silicon has a long "pipeline"), it can seem that they may be missing a new trend or swerve in the demands of the market, it is often simply because there is this delay.
- bryanlarsen 10mo agoYou could have said the same thing about Intel for ~50 years.
- tim333 10mo agoDepends on the top management though. I imagine Nvidia will keep doing well while Jensen Huang is running things.
- newyankee 10mo agoFair, but the 75% margins can be reduced to 25% with healthy competition. The lack of competition in the frontier chips space was always the bottleneck to commoditization of computation, if such a thing is even possible
- plantain 10mo agoTurkeys bet on tomorrow 364 days of the year.
- adzm 10mo agoI told you a thousand times, you have to sell your pumpkin stock before Halloween, before!
- siliconc0w 10mo agoGoogle has always had great tech - their problem is the product or the perseverance, conviction, and taste needed to make things people want.
- thomascgalvin 10mo agoTheir incentive structure doesn't lead to longevity. Nobody gets promoted for keeping a product alive, they get promoted for shipping something new. That's why we're on version 37 of whatever their chat client is called now. I think we can be reasonably sure that search, Gmail, and some flavor of AI will live on, but other than that, Google apps are basically end-of-life at launch.
- nostrademons 10mo agoIt's telling that basically all of Google's successful projects were either acquisitions or were sponsored directly by the founders (or sometimes, were acquisitions that were directly sponsored by the founders). Those are the only situations where you are immune from the performance review & promotion process.
- sidibe 10mo agoThey've actually had many very successful projects that make the few products and acquisitions you are thinking of work. It's true most of their end products don't work or get abandoned but it stretches their infrastructure in ways that works out well in the long run
- nostrademons 10mo agoI should probably have said "products" rather than "projects". There's a fair bit of extremely good engineering that goes on in the infrastructure side, but when it comes to consumer products, if one of the founders isn't explicitly sponsoring it it gets killed.
- 10mo ago
- villgax 10mo agohttps://killedbygoogle.com https://killedbygoogle.com
- riku_iki 10mo agoIt's all small products which didn't receive traction.
- davidmurdoch 10mo agoIt's not though. Chromecast, g suite legacy, podcast, music, url shortener,... These weren't small products.
- riku_iki 10mo agochromecast is alive, podcast, music were migrated to youtube app, url shortener is not core business and just side hustle for google. Not familiar with g suite legacy.
- IncreasePosts 10mo agoChromecast is "gone" because it bridged the gap of dumb tvs needing streaming capabilities. Now almost every tv sold has some kind of smart feature or can stream natively so Chromecast aren't needed.
- bgwalter 10mo agoGoogle Hangouts wasn't small. Google+ was big and supposedly "the future" and is the canonical example of a huge misallocation of resources. Google will have no problem discontinuing Google "AI" if they finally notice that people want a computer to shut up rather than talk at them.
- riku_iki 10mo ago> Google+ was big how you define big? My understanding they failed to compete with facebook, and decided to redirect resources somewhere else.
- qwertox 10mo agoHow high are the chances that as soon as China produces their own competitive TPU/GPU, they'll invade Taiwan in order to starve the West in regards to processing power, while at the same time getting an exclusive grip on the Taiwanese Fabs?
- gostsamo 10mo agoNot very. Those fabs are vulnerable things, shame if something happens to them. If China attacks, it would be for various other reasons and processors are only one of many considerations, no matter how improbable it might sound to an HN-er.
- qwertox 10mo agoWhat if China becomes self-sufficient enough to no longer rely on Taiwanese Fabs, and hence having no issues with those Fabs getting destroyed. That would put China as the leader once and for all.
- gostsamo 10mo agoFirst, the US has advanced fab capabilities and in case of a need can develop them further. On the other side, China will suffer a Russia style blockback while caught up in a nasty war with Taiwan. Totally possible, but the second order effects are much more complex than "leader once for all". The path for victory for China is not war despite the west, but a war when the west would not care.
- ricardo81 10mo agoIt's a cool subject and article and things I only have a general understanding of (considering the place of posting). What I'm sure about is having a programming unit more purposed to a task is more optimal than a general programming unit designed to accommodate all programming tasks. More and more of the economics of programming boils down to energy usage and invariably towards physical rules, the efficiency of the process has the benefit of less energy consumed. As a Layman is makes general sense. Maybe a future where productivity is based closer on energy efficiency rather than monetary gain pushes the economy in better directions. Cryptocurrency and LLMs seem like they'll play out that story over the next 10 years.
- zenoprax 10mo agoI have read in the past that ASICs for LLMs are not as simple a solution compared to cryptocurrency. In order to design and build the ASIC you need to commit to a specific architecture: a hashing algorithm for a cryptocurrency is fixed but the LLMs are always changing. Am I misunderstanding "TPU" in the context of the article?
- p-e-w 10mo agoIt’s true that architectures change, but they are built from common components. The most important of those is matrix multiplication, using a relatively small set of floating point data types. A device that accelerates those operations is, effectively, an ASIC for LLMs.
- bfrog 10mo agoWe used to call these things DSPs
- tuhgdetzhh 10mo agoWhat is the difference between a DSP and Asic? Is a GPU a DSP?
- imtringued 10mo agoA DSP contains analog to digital and digital to analog converters plus DMA for fast transfers to main memory and fixed function blocks for finite impulse response and infinite pulse response filters. The fact that they also support vector operations or matrix multiplication is kind of irrelevant and not a defining characteristic of DSPs. If you want to go that far, then everything is a DSP, because all signals are analog.
- bfrog 10mo agoSee here https://intel.github.io/intel-npu-acceleration-library/npu.html https://intel.github.io/intel-npu-acceleration-library/npu.h... Maybe also note that Qualcomm has renamed their Hexagon DSP to Hexagon NN. Likely the change was adding activation functions but otherwise its a VLIW architecture with accelerated MAC operations, aka a DSP architecture.
- paulmist 10mo ago> The GPUs were designed for graphics [...] However, because they are designed to handle everything from video game textures to scientific simulations, they carry “architectural baggage.” [...] A TPU, on the other hand, strips away all that baggage. It has no hardware for rasterization or texture mapping. With simulations becoming key to training models doesn't this seem like a huge problem for Google?
- m4r1k 10mo agoGoogle's real moat isn't the TPU silicon itself—it's not about cooling, individual performance, or hyper-specialization—but rather the massive parallel scale enabled by their OCS interconnects. To quote The Next Platform: "An Ironwood cluster linked with Google’s absolutely unique optical circuit switch interconnect can bring to bear 9,216 Ironwood TPUs with a combined 1.77 PB of HBM memory... This makes a rackscale Nvidia system based on 144 “Blackwell” GPU chiplets with an aggregate of 20.7 TB of HBM memory look like a joke." Nvidia may have the superior architecture at the single-chip level, but for large-scale distributed training (and inference) they currently have nothing that rivals Google's optical switching scalability.
- villgax 10mo ago100 times more chips for equivalent memory, sure.
- NaomiLehman 10mo agoI think it's not about the cost but the limits of quickly accessible RAM
- croon 10mo agoIronwood is 192GB, Blackwell is 96GB, right? Or am i missing something?
- villgax 10mo ago182GB and B300 is 288GB. IIRC
- m4r1k 10mo agoCheck the specs again. Per chip, TPU 7x has 192GB of HBM3e, whereas the NVIDIA B200 has 186GB. While the B200 wins on raw FP8 throughput (~9000 vs 4614 TFLOPs), that makes sense given NVIDIA has optimized for the single-chip game for over 20 years. But the bottleneck here isn't the chip—it's the domain size. NVIDIA's top-tier NVL72 tops out at an NVLink domain of 72 Blackwell GPUs. Meanwhile, Google is connecting 9216 chips at 9.6Tbps to deliver nearly 43 ExaFlops. NVIDIA has the ecosystem (CUDA, community, etc.), but until they can match that interconnect scale, they simply don't compete in this weight class.
- jimbohn 10mo agoGiven the importance of scale for this particular product, any company placing itself on "just" one layer of the whole story is at a heavy disadvantage, I guess. I'd rather have a winning google than openai or meta anyway.
- subroutine 10mo ago> I'd rather have a winning google than openai or meta anyway. Why? To me, it seems better for the market, if the best models and the best hardware were not controlled by the same company.
- jimbohn 10mo agoI agree, it would be the best of bad cases, in a sense. I have low trust in OpenAI due to its leadership, and in Meta, because, well, Meta has history, let's say.
- uselesswords 10mo agoI think you are disagreeing.
- jimbohn 10mo agoYep :)
- mosura 10mo agoThis is the “Microsoft will dominate the Internet” stage. The truth is the LLM boom has opened the first major crack in Google as the front page of the web (the biggest since Facebook), in the same way the web in the long run made Windows so irrelevant Microsoft seemingly don’t care about it at all.
- villgax 10mo agoExactly, ChatGPT pretty much ate away ad volume & retention if th already garbage search results weren't enough. Don't even get me started on Android & Android TV as an ecosystem.
- IncreasePosts 10mo agoThat's not the story that GOOGs quarterly earning reports tell(ad revenue up 12% YoY)
- pzo 10mo agomost likely because they got more aggressive with campaign against adblock in chrome and more ads in youtube.
- thesz 10mo ago5 days ago: https://news.ycombinator.com/item?id=45926371 https://news.ycombinator.com/item?id=45926371 Sparse models have same quality of results but have less coefficients to process, in case described in the link above sixteen (16) times as less. This means that these models need 8 times less data to store, can be 16 and more times faster and use 16+ times less energy. TPUs are not all that good in the case of sparse matrices. They can be used to train dense versions, but inference efficiency with sparse matrices may be not all that great.
- HarHarVeryFunny 10mo agoTPUs do include dedicated hardware, SparseCores, for sparse operations. https://docs.cloud.google.com/tpu/docs/system-architecture-tpu-vm https://docs.cloud.google.com/tpu/docs/system-architecture-t... https://openxla.org/xla/sparsecore https://openxla.org/xla/sparsecore
- thesz 10mo agoSparseCores appear to be block-sparse as opposed to element-sparse. They use 8- and 16-wide vectors to compute. Here's another inference-efficient architecture where TPUs are useless: https://arxiv.org/pdf/2210.08277 https://arxiv.org/pdf/2210.08277 There is no matrix-vector multiplication. Parameters are estimated using Gumbel-Softmax. TPUs are of no use here. Inference is done bit-wise and most efficient inference is done after application of boolean logic simplification algorithms (ABC or mockturtle). In my (not so) humble opinion, TPUs are example case of premature optimization.
- HarHarVeryFunny 10mo agoThey are on their 7th generation now, so presumably the architecture is being updated as needs require.
- 1980phipsi 10mo ago> It is also important to note that, until recently, the GenAI industry’s focus has largely been on training workloads. In training workloads, CUDA is very important, but when it comes to inference, even reasoning inference, CUDA is not that important, so the chances of expanding the TPU footprint in inference are much higher than those in training (although TPUs do really well in training as well – Gemini 3 the prime example). Does anyone have a sense of why CUDA is more important for training than inference?
- johnebgd 10mo agoI think it’s the same reason windows is inportant to desktop computers. Software was written to depend on it. Same with most of the software out there today to train being built around CUDA. Even a version difference of CUDA can break things.
- NaomiLehman 10mo agoinference is often a static, bounded problem solvable by generic compilers. training requires the mature ecosystem and numerical stability of cuda to handle mixed-precision operations. unless you rewrite the software from the ground up like Google but for most companies it's cheaper and faster to buy NVIDIA hardware
- never_inline 10mo ago> static, bounded problem What does it even mean in neural net context? > numerical stability also nice to expand a bit.
- baby_souffle 10mo agoThat quote left me with the same question. Something about decent amount of ram on one board perhaps? That’s advantageous for training but less so for inference?
- llm_nerd 10mo agoIt's just more common as a legacy artifact from when nvidia was basically the only option available. Many shops are designing models and functions, and then training and iterating on nvidia hardware, but once you have a trained model it's largely fungible. See how Anthropic moved their models from nvidia hardware to Inferentia to XLA on Google TPUs. Further it's worth noting that the Ironwood, Google's v7 TPU, supports only up to BF16 (a 16-bit floating point that has the range of FP32 minus the precision. Many training processes rely upon larger types, quantizing later, so this breaks a lot of assumptions. Yet Google surprised and actually training Gemini 3 with just that type, so I think a lot of people are reconsidering assumptions.
- jmward01 10mo agoHow much of current GPU and TPU design is based around attn's bandwith hungry design? The article makes it seem like TPUs aren't very flexible so big model architecture changes, like new architectures that don't use attn, may lead to useless chips. That being said, I think it is great that we have some major competing architectures out there. GPUs, TPUs and UMA CPUs are all attacking the ecosystem in different ways which is what we need right now. Diversity in all things is always the right answer.
- thelastgallon 10mo agoWith its AI offerings, can Google suck the oxygen out of AWS? AWS grew big because of compute. The AI spend will be far larger than compute. Can Google launch AI/Cloud offerings with free compute bundled? Use our AI, and we'll throw in compute for free.
- blinding-streak 10mo agoInteresting thought. If that strategy worked too well, I could see the government going after them for monopolistic "bundling"
- loph 10mo agoThis is highly relevant: "Meta in talks to spend billions on Google's chips, The Information reports" https://www.reuters.com/business/meta-talks-spend-billions-googles-chips-information-reports-2025-11-25/ https://www.reuters.com/business/meta-talks-spend-billions-g...
- pavelstoev 10mo agokeyword: "...talks..."
- 01100011 10mo agoWeird they'd do this after developing several generations of their own inference chip. Google is basically a competitor. This may just be a ploy to get better pricing from Nvidia.
- giardini 10mo agoAll this assumes that LLMs are the sole mechanism for AI and will remain so forever: no novel architectures (neither hardware nor software), no progress in AI theory, nothing better than LLMs, simply brute force LLM computation ad infinitum. Perhaps the assumptions are true. The mere presence of LLMs seems to have lowered the IQ of the Internet drastically, sopping up financial investors and resources that might otherwise be put to better use.
- olalonde 10mo agoThat's incorrect. TPUs can support many ML workloads, they're not exclusive to LLMs.
- thatguysaguy 10mo agoTPUs predate LLMs by a long time. They were already being used for all the other internal ML work needed for search, youtube, etc.
- kittikitti 10mo agoYou can't really buy a TPU, you have to buy the entire data center that includes the TPU plus the services and support. In Google Colab, I often don't prefer the TPU either because the documentation for the AI isn't made for it. While this could all change in the long term, I also don't see these changes in Google's long term strategy. There's also the problem with Google's graveyard which isn't mentioned in the long term of the original article. Combined with these factors, I'm still skeptical about Google's lead on AI.
- lukeschlather 10mo agoThis feels a lot like the RISC/CISC debate. More academic than it seems. Nvidia is designing their GPUs primarily to do exactly the same tasks TPUs are doing right now. Even within Google it's probably hard to tell whether or not it matters on a 5-year timeframe. It certainly gives Google an edge on some things, but in the fullness of time "GPUs" like the H100 are primarily used for running tensor models and they're going to have hardware that is ruthlessly optimized for that purpose. And outside of Google this is a very academic debate. Any efficiency gains over GPUs will primarily turn into profit for Google rather than benefit for me as a developer or user of AI systems. Since Google doesn't sell TPUs, they are extremely well-positioned to ensure no one else can profit from any advantages created by TPUs.
- turtletontine 10mo ago> Since Google doesn't sell TPUs, they are extremely well-positioned to ensure no one else can profit from any advantages created by TPUs. First part is true at the moment, not sure the second follows. Microsoft is developing their own “Maia” chips for running AI on Azure with custom hardware, and everyone else is also getting in the game of hardware accelerators. Google is certainly ahead of the curve in making full-stack hardware that’s very very specialized for machine learning. But everyone else is moving in the same direction: lots of action is in buying up other companies that make interconnects and fancy networking equipment, and AMD/NVIDIA continue to hyper specialize their data center chips for neural networks. Google is in a great position, for sure. But I don’t see how they can stop other players from converging on similar solutions.
- decimalenough 10mo agoGoogle does not sell them, but you can rent them: https://cloud.google.com/tpu https://cloud.google.com/tpu As you note, they'll set the margins to benefit themselves, but you can still eke out some benefit. Also, you can buy Edge TPUs, but as the name says these are for edge AI inference and useless for any heavy lifting workloads like training or LLMs. https://www.amazon.com/Google-Coral-Accelerator-coprocessor-Raspberry/dp/B07R53D12W https://www.amazon.com/Google-Coral-Accelerator-coprocessor-...
- torginus 10mo ago
- DonHopkins 10mo agoWill Google sell TPUs that can be plugged into stock hardware, or custom hardware with lots of TPUs? Our customers want all their video processing to happen on site, and don't want their video or other data to touch the cloud, so they're not happy about renting cloud TPUs or GPUs. Also it would be nice to have smart cameras with built-in TPUs.
- nish__ 10mo agoWhy don't your customers trust Google Cloud?
- DonHopkins 10mo agoIt's not Google Cloud per se, it's any cloud. There are a million reasons not to trust (or spend money on) any cloud. They want all their video and data on premises and completely under their control.
- a96 10mo agoWell, there's also an ever increasing list of reasons not to trust anything from Google. They're even pretty similar reasons.
- d--b 10mo agoAt this stage, it is somewhat clear that it doesn't really matter who's ahead in the race, cause everyone else is super close behind...
- Jeff-Collins 10mo ago[dead]
- hirako2000 10mo agoThen Groq should reign emperor?
- Shorel 10mo agoThey can only privatize the AI race. If Google wins, we all lose.
- wiredpancake 10mo ago[dead]
- WarOnPrivacy 10mo agoI wish we had more options for a dedicated/stand-alone TPU for end users. I recently bought a 2019 Coral, which as far as I know is my only option.
- saagarjha 10mo agoCoral really has little to do with modern TPUs.
- bastawhiz 10mo agoI don't think what the article writes about matters all that much. Gemini 3 Pro is arguably not even the best model anymore, and it's _weeks_ old, and Google has far more resources than Anthropic does. If the hardware actually was the secret sauce, Google would be wiping the floor with little everyone else. But they're not. There's a few confounding problems: 1. Actually using that hardware effectively isn't easy. It's not as simple as jacking up some constant values and reaping the benefits. Actually using the hardware is hard, and by the time you've optimized for it, you're already working on the next model. 2. This is a problem that, if you're not Google, you can just spend your way out of. A model doesn't take a petabyte of memory to train or run. Regular old H100s still mostly work fine. Faster models are nice, but Gemini 3 Pro being 50% of the latency as Opus 4.5 or GPT 5.1 doesn't add enough value to matter to really anyone. 3. There's still a lot of clever tricks that work as low hanging fruit to improve almost everything about ML models. You can make stuff remarkably good with novel research without building your own chips. 4. A surprising amount of ML model development is boots on the ground work. Doing evals. Curating datasets. Tweaking system prompts. Having your own Dyson sphere doesn't obviate a lot of the typing and staring at a screen that necessarily has to be done to make a model half decent. 5. Fancy bespoke hardware means fancy bespoke failure modes. You can search stack overflow for CUDA problems, you can't just Bing your way to victory when your fancy TPU cluster isn't doing the thing you want it to do.
- mda 10mo ago"Gemini 3 Pro is arguably not even the best model anymore" Arguably indeed, because I think it still is.
- bastawhiz 10mo agoIt definitely depends on how you're measuring. But the benchmarks don't put it at the top for many ways of measuring, and my own experience doesn't put it at the top. I'm glad if it works for you, but it's not even a month old and there are lots of folks like me who see it as definitely worse for classes of problems that 3 Pro could be the best at. Which is to say, if Google was set up to win, it shouldn't even be a question that 3 Pro is the best. It should be obvious. But it's definitely not obvious that it's the best, and many benchmarks don't support it as being the best.
- bitwize 10mo agoYes, but what's at the finish line? The bottom?
- danishSuri1994 10mo ago[dead]
- zmmmmm 10mo agoI always enjoy being wrong and I was very wrong in my predictions about Google : I thought they should theoretically win, but I was also very confident they couldn't possibly turn their execution ship around to actually pull together a coherent competitor to OpenAI. But they do seem to have done that and it's very impressive. If they do continue to execute, I can't see anybody stopping them dominating and I would be bearish on nearly every other player catching them. The biggest problem though is trust, and I'm still holding back from letting anyone under my authority in my org use Gemini because of the lack of any clear or reasonable statement or guidelines on how they use your data. I think it won't matter in the end if they execute their way to domination - but it's going to give everyone else a chance at least for a while.
- addaon 10mo ago> If they do continue to execute Yes, but Google will never be able to compete with their greatest challenge... Google's attention span.
- epistasis 10mo agoThe LLM provider I trust the most right now is AWS. Anybody else seems to have very conflicted purposes when it comes to sending them my data and interactions.
- danielheath 10mo agoYou're not wrong... but any space where Amazon, of all companies, has a shot at being the "most trustworthy player" is one I'm going to avoid where I can.
- deleted 10mo ago[deleted]
- bloppe 10mo agoAmazon makes an LLM?
- epistasis 10mo ago
- storus 10mo agoIf Google won, it would cannibalize its current ad-driven business and replace it with something that is extremely expensive to run and difficult to make profit from. A Pyrrhic win essentially.
- reval 10mo agoHardly a Pyrrhic win. When the rest of the market is burning money, whoever burns money the slowest while still remaining competitive will win.
- ezekiel68 10mo agoI mean, focus is a thing that Google has always struggled with. But I kind of doubt that customers who need online marketing (ads) are going to convert en masse to users who rent cloud TPUs instead.
- coppsilgold 10mo agoBut that would happen regardless of who won, better to at least dominate the new paradigm and figure out how to extract value from it. I also suspect that once the value generation is figured out they will cease offering these APIs to anyone, if you had a golden goose would you rent it?
- mr_toad 10mo agoThey could go all dark mirror and inject ads directly into the responses.
- anonym29 10mo agoI have never understood why, in these discussions, nobody brings up other specialized silicon providers like Groq, SambaNova, or my personal favorite, Cerebras. Cerebras CS-3 specs: • 4 trillion transistors • 900,000 AI cores • 125 petaflops of peak AI performance • 44GB on-chip SRAM • 5nm TSMC process • External memory: 1.5TB, 12TB, or 1.2PB • Trains AI models up to 24 trillion parameters • Cluster size of up to 2048 CS-3 systems • Memory B/W of 21 PB/s • Fabric B/W of 214 Pb/s (~26.75 PB/s) Comparing GPU to TPU is helpful for showcasing the advantages of the TPU in the same way that comparing CPU to Radeon GPU is helpful for showcasing the advantages of GPU, but everyone knows Radeon GPU's competition isn't CPU, it's Nvidia GPU! TPU vs GPU is new paradigm vs old paradigm. GPUs aren't going away even after they "lose" the AI inference wars, but the winner isn't necessarily guaranteed to be the new paradigm chip from the most famous company. Cerebras inference remains the fastest on the market to this day to my knowledge due to the use of massive on-chip SRAM rather than DRAM, and to my knowledge, they remain the only company focused on specialized inference hardware that has enough positive operating revenue to justify the costs from a financial perspective. I get how valuable and important Google's OCS interconnects are, not just for TPUs or inference, but really as a demonstrated PoC for computing in general. Skipping the E-O-E translation in general is huge and the entire computing hardware industry would stand to benefit from taking notes here, but that alone doesn't automatically crown Google the victor here, does it?
- sandGorgon 10mo agodeepseek kind of innovated on this using off-the-shelf components right ? to quote from their paper "In order to ensure sufficient computational performance for DualPipe, we customize efficient cross-node all-to-all communication kernels (including dispatching and combining) to conserve the number of SMs dedicated to communication. The implementation of the kernels is codesigned with the MoE gating algorithm and the network topology of our cluster."
- amitk2405 10mo agoThe part that surprised me is how much TPUs gain from the systolic array design. It basically cuts down the constant memory shuffling that GPUs have to do, so more of the chip’s time is spent actually computing. The downside is the same thing that makes them fast: they’re very specialized. If your code already fits the TPU stack (JAX/TensorFlow), you get great performance per dollar. If not, the ecosystem gap and fear of lock-in make GPUs the safer default.
- wangii 10mo agoIt's all about the right mind set at the very top level. At the beginning of the PC era, nobody would bet IBM to lose. Same in the dawn of internet, all money was on MS. so it happened to Nokia and Ericsson. Google is a giant without a direction. The ads money is so good that it just doesn't have the gut to leave it on the table.
- veunes 10mo agoThe funniest thing about this story is that NVIDIA has essentially become a TPU company. Look at the Hopper and Blackwell architectures: Tensor Cores are taking up more space, the Transformer Engine has appeared, and NVLink has started to look like a supercomputer interconnect. Jensen Huang isn't stupid. He saw the threat of specialized ASICs and just built the ASIC inside the GPU. Now we have a GPU that is 80% matrix multiplier but still keeps CUDA compatibility. Google tried to kill the GPU, but instead forced the GPU to mutate into a TPU
- breppp 10mo agoThere's an issue with building a swiss knife chip that supports everything back to the 80s, it works great until it doesn't (Intel)
- veunes 10mo agoTiming is everything here. The swiss army knife approach only loses when tasks stop changing. Intel suffered when workloads like web and mobile stabilized In AI, we're still in the explosion phase. If you build the perfect ASIC for Transformers today, and tomorrow a paper drops with a new architecture, your chip becomes a brick. NVIDIA pays the "legacy tax" and keeps CUDA specifically as insurance against algorithm churn. As long as the industry moves this fast, flexibility beats raw efficiency
- adiian 10mo ago- ASIC won the crypto mining battle in the past, it's orders of magnitude faster - Google is not owning the technology but builds a cohesive cloud around it, Tesla, Meta work on their own asic ai chips and I guess others - A signal is already given: Softbank sold it's entire Nvidia stock and berkshire added google on their portfolio. Microsoft "has" a lot of companies data, and google is probably building the most advanced ai cloud. However, I can't think they had a cloud which was light-years ahead of aws 15 years ago and now GCP is no 3, they also released opensource gpt models more than 5 years ago that constituted the foundation for openai closed sourced models.
- vagab0nd 10mo agoIf an outsider is ever allowed to have a simplified mental model of why nvidia was unbeatable, here's mine: - they were way ahead, and they didn't make any big mistakes - they weren't waiting for others to catch up. They were aggressively improving - memory bandwidth is almost always the bottleneck. Hence systolic array is "overrated". Furthermore, interconnect is the new bottleneck now - cuda offers the most flexibility in the world of ever changing model requirements
- egberts1 10mo agoA Tensor-based PCI adapter card? TAKE MY MONEY!!!