5 ms·
The fact that they didn't do this: > Their current training cluster would be the 5th largest supercomputer if Tesla stopped all real workloads, ran Linpack, an
by volta83 5y ago
The fact that they didn't do this:
> Their current training cluster would be the 5th largest supercomputer if Tesla stopped all real workloads, ran Linpack, and submitted it to the Top500 list.
which is trivial to do, and pretty much a must when bringing up the cluster to make sure its working properly, so much that most clusters do this on every maintainance, along with another bunch of benchmarks;
and that they say this:
> cost equivalent versus Nvidia GPU, Tesla claims they can achieve 4x the performance, 1.3x higher performance per watt, and 5x smaller footprint.
but have no MLPerf results, tells you everything you need to know about it.
The list of long-term hype-only AI-hardware companies with billions of dollars of VC investment and literally nothing to show is incredible and keeps growing.
Every MLPerf round, the list of companies that want to submit is "huge", and 1 week before the deadline, 99.999% of them have been saying "we'll submit next round" for years.
It's as-if people would spend billions on creating an F1 team, and then notice during pre-season training that the car can't even finish a lap. And then fail to even start a lap on every race of the season. And then do this again, year after year, for a decade. Burning billions and billions...
- formerly_proven 5y agoWhat's often overlooked is just because you have a shit-ton of compute nodes doesn't mean you could make it to the TOP500. You might have the compute power, but the system most likely doesn't have the connectivity. E.g. on the first slide it says this is distributed over more than three locations, which essentially guarantees that the system doesn't have supercomputer-like connectivity as a whole. Worth pointing out that the networking in a supercomputer is a very significant chunk of cost and power. And if it is tailored for AI, it might not even do 32-bit float, or only at a fraction of the "AI FLOPS".
- chrisseaton 5y agoI don't get it - they're using it for actual work, rather than burning power to run useless benchmarks for bragging rights - and you think that makes it hype?! Surely it's the opposite - running benchmarks rather than doing something useful is hype.
- volta83 5y agoYou have never built a super computer, have you? You have to connect thousands of cables, have hundreds of nodes, with thousands of components, everything interconnected, and if you connect one wrong, the computer outputs incorrect results. You have to routinely update the software, and if a software upgrade introduces a 20% perf regression (which happens), then your 10 MWh cluster starts burning 2MWh for nothing. Or maybe your cooling system sucks, and after a minute of running at full capacity, you need to throttle your cluster to 0.1% of the peak to keep it cool enough that it runs "something". That's why all systems in the top500 i've been involved with (15 or so) run these benchmarks as an integration tests on every single cluster maintenance (node updates, servicing, OS updates, etc.). Submitting these results to the Top500 costs you nothing... if your cluster actually works. When you submit to the Top500, they ask for access so that they can re-run them themselves, which happens typically during / after the next maintenance to avoid impacting any users. If they haven't submitted, 100% sure their cluster does not deliver what they say it should deliver on paper. Maybe it delivers 1% of it, or 0.01% of it (seen both cases in real life). If they haven't fixed it, then maybe it can't be fixed. HPL, MLPerf, Spec, Stream, OSU.... these are not "benchmarks for bragging rights", these are tests that show that your system works.
- chrisseaton 5y ago> run these benchmarks as an integration tests on every single cluster maintenance Why run someone else's benchmark and not your own application to test performance? And what's the point of submitting to Top500? Why do you care how your system ranks? What's the business or technical purpose in that?
- cedilla 5y agoFor the same reason that we don't test engines during a race. Only what's under test should change, with the rest being fixed. It would be much more difficult to adapt some of your own applications as a test. If the result are bad, how would you even know if it's the application's fault or a problem with the cluster? LINPACK on the other hand is well understood and there has been tons of work to make sure that it uses all the power your cluster can deliver.
- _ph_ 5y agoThere might be some level of showmanship involved, but Tesla isn't selling those things, it is using them. Quite a different situation in comparison to a random startup which tries to convince investors to finance their products.
- matmatmatmat 5y agoI don't know, given Elon's showmanship, is it really that different?
- elurg 5y agoIn the presentation they were quite open that they got to the point of running real loads but only on a single tile on a bench. Not clear what you're disputing here?
- volta83 5y ago> Not clear what you're disputing here? The article, that assumes "FLOPS on paper == FLOPS in practice".
- formerly_proven 5y agoI only work on a lowly private cluster but running standard benchmarks is utterly routine here (it is in fact automated). As others with HPC experience pointed out, running benchmarks is pretty much mandatory when bringing a new system up, not just to ensure it actually performs as promised, but also to weed out bad and marginal components. You do one or two weeks of intense benchmarking and testing and you're assured numerous nodes will fail and need parts replaced. When you apply patches, you benchmark again. Why? Because the suppliers fix for "code XYZ crashes nodes" is probably "let's uh just reduce power limits by a few %". Or because of unintentional issues limiting performance. When a node crashes, you benchmark it again. Why? Because linpack and friends are good at making marginal hardware fail. So it is literally unbelievable that Tesla not just stands up a cluster, but created their own hardware to do so, and didn't run any quotable benchmark and only has the theoretic FLOPS numbers for marketing.
- cogman10 5y ago> So it is literally unbelievable that Tesla not just stands up a cluster, but created their own hardware to do so, and didn't run any quotable benchmark and only has the theoretic FLOPS numbers for marketing. They haven't gotten to this point yet. They have a single tile (A node). It wouldn't surprise me if the tile they have is just a prototype as well. You have to have a cluster before you can start running cluster benchmarks.
- deleted 5y ago[deleted]
- ilovefood 5y agoThe more 3D renders in a presentation the more skepticism I develop. I noticed there is a direct leading indicator of a stock's price in relation to the quality of the 3D renders in the company's presentations/PR events. This is of course only anecdotal evidence but I'm pretty sure the hypothesis can hold it's ground against 50% of "ML Papers" today.
- fit2rule 5y ago>trivial to do >stopped all real workloads Says you. Tesla don't need to be in the Top500 list to have a supercomputer that they are using for other more useful things.
- rajnathani 5y ago> Their current training cluster would be the 5th largest supercomputer if Tesla stopped all real workloads, ran Linpack, and submitted it to the Top500 list. That line is referring to their current Nvidia A100 powered supercomputer which they set up with 5,760 A100 GPUs recently this year [0]. Read the previous line of the line which you’ve shared from the post: > Tesla has been expanding the size of their GPU clusters for years. Their current training cluster would be the 5th largest supercomputer if Tesla stopped all real workloads, ran Linpack, and submitted it to the Top500 list. [0] https://blogs.nvidia.com/blog/2021/06/22/tesla-av-training-supercomputer-nvidia-a100-gpus/ https://blogs.nvidia.com/blog/2021/06/22/tesla-av-training-s...
- clomond 5y agoHad to look into MLPerf as I don’t follow super compute. But I don’t see them not doing that test as an issue as they made it quite clear that their entire system is tailor made, like an ASIC, to focus on neural nets and specific, relevant compute pipelines to what they care about. I would imagine dropping a generic, broad based ML benchmarking tool will not only perform suboptimally but also not be representative of what they’re trying to do. It’s not meant to be a general purpose ML super computer, it’s supposed to be a super computer to solve a narrowish niche of problems.