3 ms·
Yes H100s are very expensive (think ~$40k per unit), and very hard to obtain. For this reason they're basically all sent to cloud operators who can sign large b
by nvm0n2 3y ago
Yes H100s are very expensive (think ~$40k per unit), and very hard to obtain. For this reason they're basically all sent to cloud operators who can sign large bulk deals up front and then rent them out. At that point rubber really hits the road because it's where the hardware is basically auctioned off on an hour by hour basis, so you can easily pay $40k per month for access to the cards because there's so much money flooding into AI right now.
This is why OpenAI seem to have burned through 10s of billions of dollars of Azure compute credits in, like, a few years.
Yes CUDA is basically a giant developer platform made by NVIDIA for their cards. It's libraries, tools, languages, tutorials, and other ecosystem stuff like the fact that lots of people use it. Intel and AMD might or might not need their own answer to it. They need at least parts of an answer but CUDA is a lot more than AI, so if that's all you care about, you can skip parts of it.
Now AMD has been trying to develop a CUDA competitor for a long time but it's clear either their heart isn't in it, or they just really struggle to attract skilled dev plaform people. ROCm seems to suck so hard people would rather pay NVIDIA's monopoly pricing than deal with it, which is impressively bad.
Of course the reason NVIDIA is in this position is that they supported all kinds of R&D with their CUDA effort, and it turned out to be AI that hit big. If you just implements the bits of CUDA that seem most useful you'll always be behind.