9 ms·
Tesla turns on 10k-node Nvidia H100 Cluster
- Kevcmk 3y agoNvidia is powering a mega Tesla supercomputer powered by 10,000 H100 GPUs
- ComputerGuru 3y agoDid you just repeat the headline?
- deleted 3y ago[deleted]
- jbverschoor 3y agoSo THAT's why my power blipped
- toomuchtodo 3y agoThe Dojo is open.
- kranke155 3y agoI thought Dojo was custom chips.
- ComputerGuru 3y agoYou are correct; it is, and flippant HN comments that are additionally incorrect are starting to become a thing. See the original tweet: https://twitter.com/SawyerMerritt/status/1696011140508045660 https://twitter.com/SawyerMerritt/status/1696011140508045660
- toomuchtodo 3y agoYou’re being pedantic (rightfully so) and I’m being loose with words. While Dojo is the supercomputer Tesla built for vision training, I lumped anything contributing to their machine vision model training as Dojo. It’s called Dojo because that’s where the training takes place. https://en.wikipedia.org/wiki/Tesla_Dojo https://en.wikipedia.org/wiki/Tesla_Dojo From the History section (although Technical Architecure is also worthy of consuming in its entirety): > In August 2023, Tesla powered on Dojo for production use as well as a new training cluster configured with 10,000 Nvidia H100 GPUs. I’ll take the L wrt being flippant if we’re using words very specifically in this context, that’s fair. It’s great to see Tesla expand its training resources is my sentiment, regardless of how their aggregate ML compute is segmented.
- ra7 3y agoYou’re just using it wrong. Dojo “supercomputer” specifically includes custom chips, which doesn’t exist yet.
- tomaytotomato 3y agoPain does not exist in this Dojo. Kiai!
- alecco 3y agoOriginal tweet: https://twitter.com/SawyerMerritt/status/1696011140508045660 https://twitter.com/SawyerMerritt/status/1696011140508045660 Previus article: https://www.tomshardware.com/news/teslas-dollar300-million-ai-cluster-is-going-live-today https://www.tomshardware.com/news/teslas-dollar300-million-a... This is second-hand blogspam.
- TonyTrapp 3y agoTom's Hardware and Tech Radar belong to the same company. If you consider this to be blog spam, almost any news website these days would be blog spam.
- einpoklum 3y agoAnd the original tweet is very much kool-aid heavy, with "20x performance", "30x performance" claims about the switch from one card to the next.
- 7e 3y agoOnly 10K?
- mousetree 3y ago$300 million for those 10,000
- latchkey 3y agomuch. much. more. You're not factoring in the disks, chassis, ram, networking gear, cabling, data center build, setup, install, etc etc etc...
- iamgopal 3y agoIt’s bottleneck on Nvidia side. They are producing less than Tesla consume. Tesla compute power will outclass many cloud provider combined in just three or four years with their own custom chip.
- chollida1 3y ago> It’s bottleneck on Nvidia side. They are producing less than Tesla consume. Tesla compute power will outclass many cloud provider combined in just three or four years with their own custom chip. That seems like a bold claim. Google, Microsoft and Meta make so much more money than Telsa that if making AI chips was so easy, then they could clearly out design and build Tesla without thinking too hard about it. What makes you think that Telsa, a company with far less AI workers and knowledge, and far less money than the above companies can out design and out build them?
- md_ 3y ago> What makes you think that Telsa, a company with far less AI workers and knowledge, an far less money than the above companies can out design and out build them? Presumably because Elon himself will be involved in the design, and Elon, as we all know, is one of the world's great thinkers. ;)
- kcb 3y ago> The firm also built a compute cluster fitted with 5,760 Nvidia A100 GPUs in June 2012 Wow, that's some really early hardware access. /s
- visarga 3y agoI only had about 3 NVIDIA H100 in 1980
- kcb 3y agoSomeone needs to figure out at what point all the compute in the world became more powerful than a single H100.
- Geee 3y agoLol, I was wondering if A100 is really that old. Turns out A100 was released in 2020.
- kcb 3y agoYea I assume they meant 2021. 2012 was still the early days of GPU compute. Best we had were M2090s.
- blackoil 3y agoMaybe, they picked up date when Elon first communicated that they are "ready" to go live. Like everything else it took a decade to materialize.
- jdiez17 3y agoWhat happened to their custom hardware training stack Dojo? They had some interesting ideas there. The last I heard, they had one of those tiles "working" in the lab. Pretty far from a production setup. I can imagine they either underestimated the software effort needed to squeeze as much performance as possible out of those things, or they underestimated the pace at which Nvidia scales FLOPS/$, or both.
- deleted 3y ago[deleted]
- martin8412 3y agoVaporware, just like much of what Musk talks about.
- s1gnp0st 3y agoReusable rockets, electric cars, solar panels... What would you say grants you the standing to opine here?
- mmcwilliams 3y agoI'm fairly certain all of those existed prior to Musk's suggestion of them.
- chollida1 3y agoI understand that the H100 is NVidia's leading edge chip, but can someone let me know if 10K is considered to be a big cluster? I've never worked inside one of the leading edge AI companies like OpenAI, Google, Microsoft or Meta. Is this comparable to what they would work with? My first guess is that it seems much smaller. And if you are running many parallel training jobs then you are getting about 1,000 chips at most to work with. Or is this about what the leading competitors are working with? Azure, for one, seems to have orders of magnitude more chips at their disposal.
- Proven 3y ago[dead]
- kcb 3y agoThe most powerful listed supercomputer has 37,888 Radeon GPUs, so in the same order of magnitude.
- jbverschoor 3y agoInteresting choice of words... I take you work for OpenAI? :) How large is their/'your' cluster? Probably the biggest in the world by now..
- kkielhofner 3y agoParent is almost certainly talking about Frontier, the supercomputer with the US Department of Energy[0]. [0] - https://top500.org/system/180047/ https://top500.org/system/180047/
- jbverschoor 3y agoYes, that's "listed".. I'm curious how big the "unlisted" cluster is.
- kcb 3y agoUnfortunately no, but there are almost certainly clusters in the hands of private companies and government organizations that would prefer not to advertise their capabilities.
- dahart 3y ago> This AI cluster, worth more than $300 million, will offer a peak performance of 340 FP64 PFLOPS for technical computing and 39.58 INT8 ExaFLOPS for AI applications, according to Tom’s Hardware. I was curious why this statement lead with fp64 flops (instead of fp32, perhaps), but I looked up the H100 specs, and NV’s marketing page does the same thing. They’re obviously talking about the H100 SXM here, which has the same peak theoretical fp64 throughput as fp32. The cluster perf is estimated by multiplying the GPU perf by 10k. Also, obviously, int8 tensor ops aren’t ‘FLOPS’. I think Nvidia calls them “TOPS” (tensor ops). There is a separate metric for ‘tensor flops’ or TF32.
- petermcneeley 3y agoNit: INT8 is not a floating point operation and thus cannot be used in the term "ExaFLOPS"
- queuebert 3y agoIn the old days, depending on architecture, fp64 performance could be atrocious even when fp32 was decent, so bragging about fp64 performance has an authenticity to it. Not all scientific computing requires 64 bits, but knowing that you can drop to high precision when necessary without penalty is nice. Also, back in the day, integer ops were just called 'ops', grumble grumble. But yeah FLOPS specifically refers to floating point. Calling them TOPS doesn't make sense to me, since tensor cores were meant for matrix operation speedup, and these matrices are rarely integer.
- dahart 3y agoStill true that fp64 throughput is lower for consumer GPUs - both NV and AMD. That’s kinda why I was curious about leading with that metric - outside of HPC and scientific applications, a lot of people don’t really need fp64, and the machine might normally have a much higher fp32 throughput. > knowing you can drop to high precision when necessary without penalty is nice. I guess I maybe don’t know why you’d ever have 1:1 fp32 and fp64 perf. Aren’t the fp64 multipliers (for example) basically 4x fp32 multipliers? I am under the possibly naive impression that if you have all the transistors for 1 fp64 core, that you’d end up with all the transistors you need for 2 or 4 fp32 cores. Maybe that’s not true today, but there does have to be at least 2x the transistors overall for 64-bit vs 32-bit, and lots of those should be shared or reusable, no? It doesn’t seem quite right to frame naturally higher 32-bit op throughput as a “penalty” on 64-bit ops. You’re asking the hardware to do more with 64, and it makes complete sense that given the exact same budget for bandwidth, energy, memory, compute, etc. that 32-bit ops would go faster, no? If the op throughput of fp64 and fp32 is the same, doesn’t that possibly imply that the fp32 ops are potentially being wasted / penalized, just for the sake of having matching numbers?
- bluelightning2k 3y agoIt's funny - I'm listening to "The Founders" audiobook and right now they're telling the story of Elon Musk at PayPal wanting to rewrite for Windows server because Linux was too hard. Weird to think that his next company's compute platform is this.
- WendyTheWillow 3y agoLinux was a lot harder back then.
- cactusplant7374 3y agoHarder for who? Elon certainly didn't have the technical chops to work with it.
- WendyTheWillow 3y agoHarder for everyone, including his staff, who were asking him to move to Windows…
- huggingmouth 3y agoHe should have hired staff that is competent with the tech stack used at his company. Unforced rewrites are usually always a bad idea.
- WendyTheWillow 3y agoOh that simple, huh? Too bad you weren’t around in the late 90s to explain this to him, and to help him find the extremely rare group of folks familiar with Linux…
- bluelightning2k 3y agoActually he was the other way around. He strongly pushing windows and the CTO and engineers from Confonity strongly wanted Linux
- md_ 3y agoI'm confused. The article from September 1 linked to here is strangely future-tense ("But the firm’s latest investment in 10,000 of the company’s H100 GPUs dwarfs the power of this supercomputer....This AI cluster, worth more than $300 million, will offer a peak performance..."). It links to a Tom's Hardware article (https://www.tomshardware.com/news/teslas-dollar300-million-ai-cluster-is-going-live-today https://www.tomshardware.com/news/teslas-dollar300-million-a...) from August 28 that says "Tesla is about to flip the switch on its new AI cluster, featuring 10,000 Nvidia H100 compute GPUs") and says "Tesla is set to launch its highly-anticipated supercomputer on Monday..." (presumably the September 1 event). So, like, does Tesla actually have 10k H100s? Or do they have an order for 10k H100s? Or an intention to buy 10k H100s? Is the sole source for these articles this (https://twitter.com/SawyerMerritt/status/1696011140508045660 https://twitter.com/SawyerMerritt/status/1696011140508045660) random Twitter post by some guy who runs an online clothing company? I don't mean to snipe, but this article doesn't seem to rise to the extremely high editorial standards of such tech-press luminaries as "TechRadar" and "Hacker News".
- deleted 3y ago[deleted]
- xedeon 3y ago> high editorial standards of such tech-press luminaries as "TechRadar" and "Hacker News". If you would’ve just scrolled just a little bit on that Twitter post that you linked. You would’ve seen these: https://x.com/sawyermerritt/status/1696012091964915744 https://x.com/sawyermerritt/status/1696012091964915744 https://x.com/tim_zaman/status/1695488119729238147 https://x.com/tim_zaman/status/1695488119729238147 Also, just FYI. Sawyer posts most of the Tesla and SpaceX breaking news on Twitter before major outlets even write their articles. For example, here’s one just 12mins ago as confirmed by Elon: https://x.com/sawyermerritt/status/1728092021628313777 https://x.com/sawyermerritt/status/1728092021628313777 A “random Twitter post by some guy who runs an online clothing company” is definitely a wrong assumption. https://x.com/sawyermerritt/status/1709019899442479162 https://x.com/sawyermerritt/status/1709019899442479162
- md_ 3y ago
- kaycebasques 3y agon00b questions from someone just beginning to get interested in HPC I see mention of using this supercomputer for training models. Is that the only purpose? What other types of things do orgs usually do with these supercomputers? Are there any good boots-on-the-ground technical blogs that provide interesting detail on day-to-day experiences with these things?
- abatilo 3y agoAs opposed to keeping all of your servers independent of each other, super computers are used any time you want to pretend the entire computer is one computer. In other words, they're used when you want to share some kind of state across all of the computers, without the potential overhead of communicating to some other system like a database. Physics simulations and like, molecular modeling come to mind as common examples. In the case of ML training, model parameters and broadcasting the deltas that get calculated during training are that shared state.
- cactusplant7374 3y agoIs FSD really a hardware problem for them?
- dsco 3y agoNewbie question, could this cluster easily calculate the largest prime number? I've found that the largest known prime number was found back in 2018, which is a while back considering how compute has evolved.
- astrodust 3y agoFinding the largest prime is more a contest of who's willing to commit the most ridiculous amount of compute to the goal than it is a mathematical obstacle. The cost of finding the next prime is likely into the millions now.
- andrewmcwatters 3y agoCan you imagine how much power 10,000 H100s actually produces in production? I bet you'd be able to run modern games on a cluster that large at a full 60 FPS.
- amai 3y agoDo they also order a power plant for that cluster? Or how much energy does such a thing need?
- throwaway4good 3y agoI predict it will run for 5 years and then come up with the answer: FSD needs lidar.