9 ms·
Are you aware of HIP? It's officially supported and, for code that avoids obscure features of CUDA like inline PTX, it's pretty much a find-and-replace to get a
by eslaught 2y ago
Are you aware of HIP? It's officially supported and, for code that avoids obscure features of CUDA like inline PTX, it's pretty much a find-and-replace to get a working build:
https://github.com/ROCm/HIP https://github.com/ROCm/HIP
Don't believe me? Include this at the top of your CUDA code, build with hipcc, and see what happens:
https://gitlab.com/StanfordLegion/legion/-/blob/master/runtime/hip_cuda_compat/hip_cuda.h https://gitlab.com/StanfordLegion/legion/-/blob/master/runti...
It's incomplete because I'm lazy but you can see most things are just a single #ifdef away in the implementation.
- currymj 2y agoif you're talking about building anything, that is already too hard for ML researchers. you have to be able to pip install something and just have it work, reasonably fast, without crashing, and also it has to not interfere with 100 other weird poorly maintained ML library dependencies.
- bootsmann 2y agoDon’t most orgs that are deep enough to run custom cuda kernels have dedicated engineers for this stuff. I can’t imagine a person who can write raw cuda not being able to handle things more difficult than pip install.
- gaogao 2y agoEngineers who are really, really good at CUDA are worth their weight in gold, so there's more projects for them than they have time. Worth their weight in gold isn't figurative here – the one I know has a ski house more expensive than 180 lbs of gold (~$5,320,814).
- bbkane 2y agoWould you (or your friend) be able to drop any good CUDA learning resources? I'd like to be worth my weight in gold...
- throwaway81523 2y agoA working knowledge of C++, plus a bit of online reading about CUDA and the NVidia GPU architecture, plus studying the LCZero chess engine source code (the CUDA neural net part, I mean) seems like enough to get started. I did that and felt like I could contribute to that code, at least at a newbie level, given the hardware and build tools. At least in the pre-NNUE era, the code was pretty readable. I didn't pursue it though. Of course becoming "really good" is a lot different and like anything else, it presumably takes a lot of callused fingertips (from typing) to get there.
- 8n4vidtmkvmk 2y agoDoes this pay more than $500k/yr? I already know C++, could be tempted to learn CUDA.
- throwaway81523 2y agoI kinda doubt it. Nobody paid me to do that though. I was just interested in LCZero. To get that $500k/year, I think you need up to date ML understanding and not just CUDA. CUDA is just another programming language while ML is a big area of active research. You could watch some of the fast.ai ML videos and then enter some Kaggle competitions if you want to go that route.
- almostgotcaught 2y agoYou're wrong. The people building the models don't write CUDA kernels. The people optimizing the models write CUDA kernels. And you don't need to know a bunch of ML bs to optimize kernels. Source: I optimize GPU kernels. I don't make 500k but I'm not that far from.
- throwaway81523 2y agoHeh I'm in the wrong business then. Interesting. Used to be that game programmers spent lots of time optimizing non-ML CUDA code. They didn't make anything like 500k at that time. I wonder what the ML industry has done to game development, or for that matter to scientific programming. Wow.
- eigenvalue 2y agoThat’s pretty funny. Good test of value across the millennia. I wonder if the best aqueduct engineers during the peak of Ancient Rome’s power had villas worth their body weight in gold.
- Willish42 2y agoThe fact that "worth their weight in cold" typically means in the single-digit millions is fascinating to me (though I doubt I'll be able to get there myself, maybe someday). I looked it up though and I think this is undercounting the current value of gold per ounce/lb/etc. 5320814 / 180 / 16 = ~1847.5 Per https://www.apmex.com/gold-price https://www.apmex.com/gold-price and https://goldprice.org/ https://goldprice.org/, current value is north of $2400 / oz. It was around $1800 in 2020. That growth for _gold_ of all things (up 71% in the last 5 years) is crazy to me. It's worth noting that anyone with a ski house that expensive probably has a net worth well over twice the price of that ski house. I guess it's time to start learning CUDA!
- boulos 2y agoNote: gold uses troy ounces, so adjust by ~10%. It's easier to just use grams or kilograms :).
- Willish42 2y agoThanks, I'm a bit new to this entire concept. Do troy lbs also exist, or is that just a term when measuring ounces?
- someguydave 2y agoyes, there are troy pounds, they are 12 troy ounces (not 16 ounces, like normal (avoirdupois) pounds) https://en.wikipedia.org/wiki/Troy_weight https://en.wikipedia.org/wiki/Troy_weight 180 avoirdupois pounds is 2,625 ounces troy. The gold price is around $2470/ounce troy today, so $2470*2625 ~= $6.483 million
- atwrk 2y ago> That growth for _gold_ of all things (up 71% in the last 5 years) is crazy to me. For comparison: S&P500 grew about the same during that period (more than 100% from Jan 2019, about 70 from Dec 2019), so the higher price of gold did not outperform the growth of the general (financial) economy.
- iftheshoefitss 2y agoWhat do people study to figure out CUDA? I’m studying to get me GED and hope to go to school one day
- paulmd 2y agoComputer science. This is a grad level topic probably. Nvidia literally wrote most of the textbooks in this field and you’d probably be taught using one of these anyway: https://developer.nvidia.com/cuda-books-archive https://developer.nvidia.com/cuda-books-archive “GPGPU Gems” is another “cookbook” sort of textbook that might be helpful starting out but you’ll want a good understanding of the SIMT model etc.
- amelius 2y agoJust wait until someone trains an ML model that can translate any CUDA code into something more portable like HIP. GP says it is just some #ifdefs in most cases, so an LLM should be able to do it, right?
- FuriouslyAdrift 2y agoOpenAI Triton? Pytorch 2.0 already uses it. https://openai.com/index/triton/ https://openai.com/index/triton/
- varelse 2y ago[dead]
- radarsat1 2y agoSelection bias. I'm sure there are lots of people who are really good at CUDA and don't have those kind of assets. Not everyone knows how to sell their skills.
- smallnamespace 2y agoUnfortunately it's also hard to buy (find) people who don't know how to sell.
- Der_Einzige 2y agoRight now, nvidias valuations have made a lot of people realize that their CUDA skills were being undervalued. Anyone with GPU or ML skills who hasn’t tried to get a pay raise in this market deserves exactly the life that they are living.
- phkahler 2y ago>> Don’t most orgs that are deep enough to run custom cuda kernels have dedicated engineers for this stuff. I can’t imagine a person who can write raw cuda not being able to handle things more difficult than pip install. This seems to be fairly common problem with software. The people who create software regularly deal with complex tool chains, dependency management, configuration files, and so on. As a result they think that if a solutions "exists" everything is fine. Need to edit a config file for your particular setup? No problem. The thing is, I have been programming stuff for decades and I really hate having to do that stuff and will avoid tools that make me do it. I have my own problems to solve, and don't want to deal with figuring out tools no matter how "simple" the author thinks that is to do. A huge part of the reason commercial software exists today is probably because open source projects don't take things to this extreme. I look at some things that qualify as products and think they're really simplistic, but they take care of some minutia that regular people are will to pay so they don't have to learn or deal with it. The same can be true for developers and ML researchers or whatever.
- jchw 2y agoThe target audience of interoperability technology is whoever is building, though. Ideally, interoperability technology can help software that supports only NVIDIA GPUs today go on to quickly add baseline support for Intel and AMD GPUs tomorrow. (and for one data point, I believe Blender is actively using HIP for AMD GPU support in Cycles.)
- Agingcoder 2y agoTheir target is hpc users, not ml researchers. I can understand why this would be valuable to this particular crowd.
- eslaught 2y agoIf your point is that HIP is not a zero-effort porting solution, that is correct. HIP is a low-effort solution, not a zero effort solution. It targets users who already use and know CUDA, and minimizes the changes that are required from pre-existing CUDA code. In the case of these abstraction layers, then it would be the responsibility of the abstraction maintainers (or AMD) to port them. Obviously, someone who does not even use CUDA would not use HIP either. To be honest, I have a hard time believing that a truly zero-effort solution exists. Especially one that gets high performance. Once you start talking about the full stack, there are too many potholes and sharp edges to believe that it will really work. So I am highly skeptical of original article. Not that I wouldn't want to be proved wrong. But what they're claiming to do is a big lift, even taking HIP as a starting point. The easiest, fastest (for end users), highest-performance solution for ML will come when the ecosystem integrates it natively. HIP would be a way to get there faster, but it will take nonzero effort from CUDA-proficient engineers to get there.
- currymj 2y agoI agree completely with your last point. As other commenters have pointed out, this is probably a good solution for HPC jobs where everyone is using C++ or Fortran anyway and you frequently write your own CUDA kernels. From time to time I run into a decision maker who understandably wants to believe that AMD cards are now "ready" to be used for deep learning, and points to things like the fact that HIP mostly works pretty well. I was kind of reacting against that.
- ezekiel68 2y ago> if you're talking about building anything, that is already too hard for ML researchers. I don't think so. I agree it is too hard for the ML researches at the companies which will have their rear ends handed to them by the other companies whose ML researchers can be bothered to follow a blog post and prompt ChatGPT to resolve error messages.
- jokethrowaway 2y agoa lot of ML researchers stay pretty high level and reinstall conda when things stop working and rightly so, they have more complicated issues to tackle It's on developers to provide better infrastructure and solve these challenges
- currymj 2y agoI'm not really talking about companies here for the most part, I'm talking about academic ML researchers (or industry researchers whose role is primarily academic-style research). In companies there is more incentive for good software engineering practices. I'm also speaking from personal experience: I once had to hand-write my own CUDA kernels (on official NVIDIA cards, not even this weird translation layer): it was useful and I figured it out, but everything was constantly breaking at first. It was a drag on productivity and more importantly, it made it too difficult for other people to run my code (which means they are less likely to cite my work).
- elashri 2y agoAs someone doing a lot of work with CUDA in a big research organization, there are few of us. If you are working with CUDA, then you are not from the type of people who wait to have something that just works like you describe. CUDA itself is a battle with poorly documented stuff.
- klik99 2y agoGod this explains so much about my last month, working with tensorflow lite and libtorch in C++
- SushiHippie 2y agoAMD has hipify for this, which converts cuda code to hip. https://github.com/ROCm/HIPIFY https://github.com/ROCm/HIPIFY
- 3abiton 2y agoThere is more glaring issue, ROCm doesn't even work well on most AMD devices nowadays, and hip performance wise deterioriates on the same hardware compared to ROCm.
- boroboro4 2y agoIt supports all of current datacenter GPUs. If you want to write very efficient CUDA kernel for modern datacenter NVIDIA GPU (read H100), you need to write it with having hardware in mind (and preferably in hands, H100 and RTX 4090 behave very differently in practice). So I don't think the difference between AMD and NVIDIA is as big as everyone perceives.
- jph00 2y agoInline PTX is hardly an obscure feature. It's pretty widely used in practice, at least in the AI space.
- saagarjha 2y agoYeah, a lot of the newer accelerators are not even available without using inline PTX assembly. Even the ones that are have weird shapes that are not amenable to high-performance work.
- HarHarVeryFunny 2y agoAre you saying that the latest NVIDIA nvcc doesn't support the latest NVIDIA devices?
- adrian_b 2y agoFor any compiler, "supporting" a certain CPU or GPU only means that they can generate correct translated code with that CPU or GPU as the execution target. It does not mean that the compiler is able to generate code that has optimal performance, when that can be achieved by using certain instructions without a direct equivalent in a high-level language. No compiler that supports the Intel-AMD ISA knows how to use all the instructions available in this ISA.
- HarHarVeryFunny 2y agoSure, but I'm not sure if that is what the parent poster was saying (that nvcc generates poor quality PTX for newer devices). It's been a while since I looked at CUDA, but it used to be that NVIDIA were continually extending cuDNN to add support for kernels needed by SOTA models, and I assume these kernels were all hand optimized. I'm curious what kind of models people are writing where not only is there is no optimized cuDNN support, but also solutions like Triton or torch.compile, and even hand optimized CUDA C kernels are too slow. Are hand written PTX kernels really that common ?
- pjmlp 2y agoHow does it run CUDA Fortran?