4 ms·
not sure why people give so much credit to Cuda, at least now. AI hasn't cared about Cuda for years - sure, Torch on NV will use it, but AI is all about Torch/
by markhahn 2y ago
not sure why people give so much credit to Cuda, at least now.
AI hasn't cared about Cuda for years - sure, Torch on NV will use it, but AI is all about Torch/TF, not details below that.
- rhaps0dy 2y agoWhat? This is really wrong. The name of the game these days is optimizing memory and FLOPs usage on these GPUs. For example, Hopper (H100) cards introduced the WGMMA (Warp-group Matrix-Multiply Accumulate) instruction, and anyone who does LLM training jumped to use it ASAP because without it you can't fully utilize the FLOPs. Anyone training AIs at the cutting age cares a lot about CUDA details. https://hazyresearch.stanford.edu/blog/2024-05-12-tk https://hazyresearch.stanford.edu/blog/2024-05-12-tk
- hnaccount_rng 2y agoThat’s not what OP meant though. That’s caring about _hardware_ capability. You could do that yourself for different hardware. And (OP’s point) any alternative hardware provider could do that for you. It just happens to be really, really hard
- rhaps0dy 2y agoI think "AI is all about Torch/TF, not details below that" directly contradicts the fact that ML people very much care about the details, to squeeze performance out of the hardware.
- nemothekid 2y ago>AI is all about Torch/TF, not details below that This could be 100% true, and the best ML engineers in the world could be completely oblivious to CUDA and it wouldn't change why nvidia is leading right now. Anyone who thinks it's "trivial" to use AMD cards should help geohot and his tinygrad boxes. Despite emailing the CEO several times, it was a clear uphill battle to get something working, and even now the sales page lists driver quality as "mediocre". If you open a GitHub issue or a forum post about your weird CUDA application that isn't working right, you will have an engineer help you. This is in nvidia's DNA and they did this before AI started printing money. On the flipside, you can email Lisa Su today and still not have problems fixed in AMDs drivers. I'm sure AMD is trying, this isn't to say they aren't, but having decent software culture at these large companies is hard. If AMD can't do it, I feel even less confident about Intel. If the CUDA moat truly didn't matter, we would see AMD cards in demand. However it doesn't and it isn't for lack of trying.
- hnaccount_rng 2y agoI literally sat at a table where the argument was made: The Intel accelerator is faster. It provides more FLOPs. But you can’t get actual users to use it outside of our LINPack runs, since it just doesn’t provide bindings to all the software actually intended to run. Granted that was in the HPC space (ie FP64 performance), but that is literally the argument. That doesn’t mean, that ML people don’t care about some hardware specifics btw. Of course they do! Because CUDA is very, very good at exploiting the specific features of the newer cards (I’ve forgotten, what exactly this was though). But take away the simple integration into their favorite AI-SDK and it would be completely irrelevant. The point is just that, if Intel were to build strictly better hardware, but not provide the interface to torch, then they would still loose to Nvidia. Yes, yes, this is not completely true. Given enough advantage someone else would build the IntelCUDA. But this is literally the bet that those academics make. And it’s so risky, that only academics are doing that
- hnaccount_rng 2y agoBecause Torch only exists on CUDA-platforms. At least in any useful form. This was, if anything, the main theme at Supercomputing: “there is not even a point in talking about Nvidia’s new benchmarks. You all are going to buy it anyways” coupled with very few “we are betting on buying cheap hardware (Intel/AMD GPUs) and hope that we can build the relevant parts of CUDA ourselves” and the latter is pretty much a desperation move of labs/sites that simply cannot get NVIDIA GPUs (either price or availability). And yes that is probably the seed of the end of Nvidia’s dominance. But it will take 20 years and multiple fuck ups. Just as it did with Intel
- MBCook 2y agoMaybe not now (I’m not in a position to know), but wasn’t it a HUGE factor in them becoming the player they are in non-graphics stuff? Because they had great APIs/libraries/documentation to play in that space?