20 ms·
Run CUDA, unmodified, on AMD GPUs
- dagmx 2y agoHas anyone tried this and knows how well it works? It definitely sounds very compelling
- arjvik 2y agoWho is this Spectral Compute, and where can we see more about them?
- msond 2y agoYou can learn more about us on https://spectralcompute.co.uk https://spectralcompute.co.uk
- JonChesterfield 2y agoThe branch free regex engine is an interesting idea. I would have said that can't be implemented in finite code. Compile to DFA by repeatedly differentiating then unroll the machine? You'd still have back edges for the repeating sections.
- pixelpoet 2y agoIsn't this a bit legally dubious, like zluda?
- janice1999 2y agoIt's advertised as a "clean room" re-implementation. What part would be illegal?
- ekelsen 2y agoIf they had to reverse engineer any compiled code to do this, I think that would be against licenses they had to agree to? At least grounds for suing and starting an extensive discovery process and possibly a costly injunction...
- msond 2y agoWe have not reverse engineered any compiled code in the process of developing SCALE. It was clean-room implemented purely from the API surface and by trial-and-error with open CUDA code.
- RockRobotRock 2y agoIsn't that exactly what a "clean room" approach avoids?
- ekelsen 2y agooh definitely. But if I was NVIDIA I'd want to verify that in court after discovery rather than relying on their claim on a website.
- RockRobotRock 2y agogood point
- ekelsen 2y agoFWIW, I think this is really great work and I wish only the best for scale. Super impressed.
- Keyframe 2y agoCan't run useful shit on it: https://docs.nvidia.com/deeplearning/cudnn/latest/reference/eula.html https://docs.nvidia.com/deeplearning/cudnn/latest/reference/... Namely: "4.1 License Scope. The SDK is licensed for you to develop applications only for use in systems with NVIDIA GPUs."
- mkl 2y agoSo add a cheap NVidia card alongside grunty AMD ones, and check for its existence. It doesn't seem to say it needs to run on NVidia GPUs.
- Keyframe 2y agoHeh, true. On the other hand, I bet companies are eager to challenge the wrath of a $3T company for a promise of "maybe it'll work, not all of it but at least it'll run worse, at least for now".
- JonChesterfield 2y agoI don't think the terms of the Nvidia SDK can restrict running software without said SDK. Nvidia's libraries don't seem to be involved here. Their hardware isn't involved either. It's just some ascii in a bunch of text files being hacked around with before running on someone else's hardware.
- adzm 2y agoI'd love to see some benchmarks but this is something the market has been yearning for.
- msond 2y agoWe're putting together benchmarks to publish at a later time, and we've asked some independent third parties to work on their own additionally.
- _lvbh 2y agoImpressive if true. Unfortunately not open source and scarce on exact details on how it works Edit: not sure why I just sort of expect projects to be open source or at least source available these days.
- tempaccount420 2y agoThey might be hoping to be acquired by AMD
- ipsum2 2y agoThey're using Docusaurus[1] for their website, which is most commonly used with open source projects. https://docusaurus.io/docs https://docusaurus.io/docs
- msond 2y agoActually, we use mkdocs and the excellent material for mkdocs theme: https://squidfunk.github.io/mkdocs-material/ https://squidfunk.github.io/mkdocs-material/
- msond 2y agoWe're going to be publishing more details on later blog posts and documentation about how this works and how we've built it. Yes, we're not open source, however our license is very permissive. It's both in the software distribution and viewable online at https://docs.scale-lang.com/licensing/ https://docs.scale-lang.com/licensing/
- breck 2y agoHow about trying _Early_ Source? It's open source with a long delay, but paying users get the latest updates. Make the git repo from "today - N years" open source, where N is something like 1 or 2. That way, students can learn on old versions, and when they grow into professionals they can pay for access to the cutting Edge builds. Win win win win ( https://breckyunits.com/earlySource.html https://breckyunits.com/earlySource.html)
- juujian 2y agoI don't understand how AMD has messed up so badly that I feel like celebrating a project like this. Features of my laptop are just physically there but not usable, particularly in Linux. So frustrating.
- djbusby 2y agoSame boat, AMD CPU but nothing else. I feel like a moderate improvement of their FOSS support, drivers would open new hardware revenue - to say nothing about the AI channel.
- ActorNightly 2y agoI don't know if I would call it a mess up. AMD still has massive market in server chips, and their ARM stuff is on the horizon. We all assume that graphics cards are the way forward for ML, which may not be the case in the future. Nvidia were just ahead in this particular category due to CUDA, so AMD may have just let them run with it for now.
- jeroenhd 2y agoAMD hardware works fine, the problem is that the major research projects everyone copies are all developed specifically for Nvidia. Now AMD is spinning up CUDA compatibility layer after CUDA compatibility layer. It's like trying to beat Windows by building another ReactOS/Wine. It's an approach doomed to fail unless AMD somehow manages to gain vastly more resources than the competition. Apple's NPU may not be very powerful, but many models have been altered specifically to run on them, making their NPUs vastly more useful than most equivalently powerful iGPUs. AMD doesn't have that just yet, they're always catching up. It'll be interesting to see what Qualcomm will do to get developers to make use of their NPUs on the new laptop chips.
- JonChesterfield 2y agoInteresting analogy. The last few programs from the windows world I tried to run were flawless under wine and abjectly failed under windows 11.
- deliveryboyman 2y agoWould like to see benchmarks for the applications in the test suite. E.g., how does Cycles compare on AMD vs Nvidia?
- Straw 2y agoI worked for spectral compute a few years ago. Very smart and capable technical team. At the time, not only did they target AMD (with less compatibility than they have now), but also outperformed the default LLVM ptx backend, and even NVCC, when compiling for Nvidia GPUs!
- modeless 2y agoA lot of people think AMD should support these translation layers but I think it's a bad idea. CUDA is not designed to be vendor agnostic and Nvidia can make things arbitrarily difficult both technically and legally. For example I think it would be against the license agreement of cuDNN or cuBLAS to run them on this. So those and other Nvidia libraries would become part of the API boundary that AMD would need to reimplement and support. Chasing bug-for-bug compatibility is a fool's errand. The important users of CUDA are open source. AMD can implement support directly in the upstream projects like pytorch or llama.cpp. And once support is there it can be maintained by the community.
- DeepYogurt 2y agoYa, honestly better to leave that to third parties who can dedicate themselves to it and maybe offer support or whatever. Let AMD work on good first party support first.
- fngjdflmdflg 2y ago>Nvidia can make things arbitrarily difficult both technically and legally. I disagree. AMD can simply not implement those APIs, similar to how game emulators implement the most used APIs first and sometimes never bother implementing obscure ones. It would only matter that NVIDIA added eg. patented APIs to CUDA if those APIs were useful. In which case AMD should have a way to do them anyway. Unless NVIDIA comes up with a new patented API which is both useful and impossible to implement in any other way, which would be bad for AMD in any event. On the other hand, if AMD start supporting CUDA and people start using AMD cards, then developers will be hesitant to use APIs that only work on NVIDIA cards. Right now they are losing billions of dollars on this. Then again they barely seem capable of supporting RocM on their cards, much less CUDA. You have a fair point in terms of cuDNN and cuBLAS but I don't know that that kind of ToS is actually binding.
- avidphantasm 2y agoPatented API? I thought Google v. Oracle settled this? Making an implementation of an API spec is fair use, is it not?
- jarbus 2y agoReally, really, really curious as to how they managed to pull this off, if their project works as well as they claim it does. If stuff as complex as paged/flash attention can "just work", this is really cool.
- Straw 2y agoMy understanding from chatting with them is that tensor core operations aren't supported yet, so FlashAttention likely won't work. I think its on their to-do list though! Nvidia actually has more and more capable matrix multiplication units, so even with a translation layer I wouldn't expect the same performance until AMD produces better ML cards. Additionally, these kernels usually have high sensitivity to cache and smem sizes, so they might need to be retuned.
- Der_Einzige 2y agoSo the only part that anyone actually cares about, as usual, is not supported. Same story as it was in 2012 with AMD vs Nvidia (and likely much before that too!). The more things change, the more they stay the same.
- JonChesterfield 2y agoCuda is a programming language. You implement it like any other. The docs are a bit sparse but not awful. Targeting amdgpu is probably about as difficult as targeting x64, mostly changes the compiler runtime. The online ptx implementation is notable for being even more annoying to deal with than the cuda, but it's just bytes in / different bytes out. No magic.
- m3kw9 2y agoThis isn’t a solution for pros because it will always play catch up and Nvidia can always add things to make it difficult. This is like emulation.
- bachmeier 2y ago> it will always play catch up That's not important if the goal is to run existing CUDA code on AMD GPUs. All you have to do is write portable CUDA code in the future regardless of what Nvidia does if you want to keep writing CUDA. I don't know the economics here, but if the AMD provides a significant cost saving, companies are going to make it work. > Nvidia can always add things to make it difficult Sounds like Microsoft embedding the browser in the OS. It's hard to see how doing something like that wouldn't trigger an antitrust case.
- dboreham 2y agoPros will end up overruled by bean counters if it works.
- ok123456 2y agoIt's not emulation. It's a compiler.
- joe_the_user 2y agoThis sounds fabulous. I look forward to AMD being drawn kicking and screaming into direct competition with Nvidia.
- gizajob 2y agoIs Nvidia not likely to sue or otherwise bork this into non-existence?
- chx 2y agoSue over what...?
- gizajob 2y agoWhatever IP related issues they’d want to sue over. Sorry I don’t know specifics about what this would specifically infringe but I’m sure expensive legal brains could come up with something
- CoastalCoder 2y agoI wonder if nVidia's current anti-trust woes would make them reluctant to go that route at the moment.
- sakras 2y agoOne question I always have about these sorts of translation layers is how they deal with the different warp sizes. I'd imagine a lot of CUDA code relies on 32-wide warps, while as far as I know AMD tends to have 64-wide warps. Is there some sort of emulation that needs to happen?
- mpreda 2y agoThe older AMD GCN had 64-wide wavefront, but the newer AMD GPUs "RDNA" support both 64 and 32 wavefront, and this is configurable at runtime. It appears the narrower wavefronts are better suited for games in general. Not sure what is the situation with "CDNA", which is the compute-oriented evolution of "GCN", i.e. whether CDNA is 64-wavefront only or dual like RNDA.
- msond 2y agoSCALE is not a "translation layer", it's a full source-to-target compiler from CUDA-like C++ code to AMD GPUs. See this part of the documentation for more details regarding warp sizes: https://docs.scale-lang.com/manual/language-extensions/#improved-support-for-non-32-warpsize https://docs.scale-lang.com/manual/language-extensions/#impr...
- ladberg 2y agoI don't really see how any code that depends heavily on the underlying hardware can "just work" on AMD. Most serious CUDA code is aware of register file and shared memory sizes, wgmma instructions, optimal tensor core memory & register layouts, tensor memory accelerator instructions, etc... Presumably that stuff doesn't "just work" but they don't want to mention it?
- lmeyerov 2y agoSort of A lot of our hw-aware bits are parameterized where we fill in constants based on the available hw . Doable to port, same as we do whenever new Nvidia architectures come out. But yeah, we have tricky bits that inline PTX, and.. that will be more annoying to redo.
- Retr0id 2y ago> SCALE accepts CUDA programs as-is. [...] This is true even if your program uses inline PTX asm
- lmeyerov 2y agoOh that will be interesting to understand, as PTX gets to more about trickier hw-arch-specific phenomena that diff brands disagree on, like memory models. Neat!
- lmeyerov 2y agoLooks like the PTX translation is via another project ZLUDA, though how they bridge the differences in memory/consistency/etc models safely remains unclear to me...
- ckitching 2y agoHi! Spectral engineer here! SCALE does not use any part of ZLUDA. We have modified the clang frontend to convert inline PTX asm block to LLVM IR. To put in a less compiler-engineer-ey way: for any given block of PTX, there exists a hypothetical sequence of C++/CUDA code you could have written to achieve the same effect, but on AMD (perhaps using funky __builtin_... functions if the code includes shuffles/ballots/other-weird-gpu-stuff). Our compiler effectively converts the PTX into that hypothetical C++. Regarding memory consistency etc.: NVIDIA document the "CUDA memory consistency model" extremely thoroughly, and likewise, the consistency guarantees for PTX. It is therefore sufficient to ensure that we use operations at least as synchronising as those called for in the documented semantics of the language (be it CUDA or PTX, for each operation). Differing consistency _between architectures_ is the AMDGPU backend's problem.
- shmerl 2y agoCompiler isn't open source? That feels like DOA in this day and age. There is ZLUDA already which is open. If they plan to open it up, it can be something useful to add to options of breaking CUDA lock-in.
- uyzstvqs 2y agoZLUDA is pretty good, except that it lacks cuDNN which makes most PyTorch projects just not work. Not sure if this project does cover that? That could be a game changer, otherwise yeah ZLUDA is the better open-source option.
- cheptsov 2y agoSounds really awesome. Any chance someone can suggest if this works also inside a Docker container?
- ckitching 2y agoIt works exactly as well as other AMDGPU-related software (HIP etc.) works inside Docker. There are some delightful AMD driver issues that make certain models of GPU intermittently freeze the kernel when used from docker. That was great fun when building SCALE's CI system :D.
- cheptsov 2y agoWould love to give it a try! Thanks for answering my question.
- SushiHippie 2y agoWorks like described in the rocm documentation (at least the scaleinfo worked for me, haven't tested further) https://rocm.docs.amd.com/projects/install-on-linux/en/latest/how-to/docker.html#docker-access-gpus-in-container https://rocm.docs.amd.com/projects/install-on-linux/en/lates...
- cheptsov 2y agoThank you! This link is very helpful.
- cheptsov 2y agoWow, somebody doesn’t like Docker enough to downvote my question.
- bornfreddy 2y agoHere, compensated with an upvote. It is a legit question.
- resters 2y agoThe main cause of Nvidia's crazy valuation is AMD's unwillingness to invest in making its GPUs as useful as Nvidia's for ML. Maybe AMD fears antitrust action, or maybe there is something about its underlying hardware approach that would limit competitiveness, but the company seems to have left billions of dollars on the table during the crypto mining GPU demand spike and now during the AI boom demand spike.
- karolist 2y agoI think this could be cultural differences, AMD's software department is underfunded and doing poorly for a long time now. * https://www.levels.fyi/companies/amd/salaries/software-engineer?country=254 https://www.levels.fyi/companies/amd/salaries/software-engin... * https://www.levels.fyi/companies/nvidia/salaries/software-engineer?country=254 https://www.levels.fyi/companies/nvidia/salaries/software-en... And it's probably better now. Nvidia was paying much more long before, also their stock growing attracts even more talent.
- 1024core 2y ago> I think this could be cultural differences, AMD's software department is underfunded and doing poorly for a long time now. Rumor is that ML engineers (that AMD really needs) are expensive; and AMD doesn't want to give them more money than the rest of the SWEs they have (for pissing off the existing SWEs). So AMD is caught in a bind: can't pay to get top MLE talent and can't just sit by and watch NVDA eat its lunch.
- paulmist 2y agoDoesn't seem to mention CDNA?
- JonChesterfield 2y agoThis is technically feasible so might be the real thing. Parsing inline ptx and mapping that onto amdgpu would be a huge pain. Working from cuda source that doesn't use inline ptx to target amdgpu is roughly regex find and replace to get hip, which has implemented pretty much the same functionality. Some of the details would be dubious, e.g. the atomic models probably don't match, and volta has a different instruction pointer model, but it could all be done correctly. Amd won't do this. Cuda isn't a very nice thing in general and the legal team would have kittens. But other people totally could.
- ckitching 2y ago[I work on SCALE] Mapping inline ptx to AMD machine code would indeed suck. Converting it to LLVM IR right at the start of compilation (when the initial IR is being generated) is much simpler, since it then gets "compiled forward" with the rest of the code. It's as if you wrote C++/intrinsics/whatever instead. Note that nvcc accepts a different dialect of C++ from clang (and hence hipcc), so there is in fact more that separates CUDA from hip (at the language level) than just find/replace. We discuss this a little in [the manual](https://docs.scale-lang.com/manual/dialects/ https://docs.scale-lang.com/manual/dialects/) Handling differences between the atomic models is, indeed, "fun". But since CUDA is a programming language with documented semantics for its memory consistency (and so is PTX) it is entirely possible to arrange for the compiler to "play by NVIDIA's rules".
- JonChesterfield 2y agoHuh. Inline assembly is strongly associated in my mind with writing things that can't be represented in LLVM IR, but in the specific case of PTX - you can only write things that ptxas understands, and that probably rules out wide classes of horrendous behaviour. Raw bytes being used for instructions and for data, ad hoc self modifying code and so forth. I believe nvcc is roughly an antique clang build hacked out of all recognition. I remember it rejecting templates with 'I' as the type name and working when changing to 'T', nonsense like that. The HIP language probably corresponds pretty closely to clang's cuda implementation in terms of semantics (a lot of the control flow in clang treats them identically), but I don't believe an exact match to nvcc was considered particularly necessary for the clang -x cuda work. The ptx to llvm IR approach is clever. I think upstream would be game for that, feel free to tag me on reviews if you want to get that divergence out of your local codebase.
- ur-whale 2y agoIf this actually works (remains to be seen), I can only say: 1) Kudos 2) Finally !
- gedy 2y agoor: 1) CUDAs
- anthonix1 2y agoI just tried it with llm.c ... seems to be missing quite a few key components such as cublaslt, bfloat16 support, nvtx3, compiler flags such as -t And its linked against an old release of ROCm. So unclear to me how it is supposed to be an improvement over something like hipify
- ckitching 2y agoGreetings, I work on SCALE. It appears we implemented `--threads` but not `-t` for the compiler flag. Oeps. In either case, the flag has no effect at present, since fatbinary support is still in development, and that's the only part of the process that could conceivably be parallelised. That said: clang (and hence the SCALE compiler) tends to compile CUDA much faster than nvcc does, so this lack of the parallelism feature is less problematic than it might at first seem. NVTX support (if you want more than just "no-ops to make the code compile") requires cooperation with the authors of profilers etc., which has not so far been available bfloat16 is not properly supported by AMD anyway: the hardware doesn't do it, and HIP's implementatoin just lies and does the math in `float`. For that reason we haven't prioritised putting together the API. cublasLt is a fair cop. We've got a ticket :D.
- anthonix1 2y agoHi, why do you believe that bfloat16 is not supported? Can you please provide some references (specifically the part about the hardware "doesn't do it")? For the hardware you are focussing on (gfx11), the reference manual [2] and the list of LLVM gfx11 instructions supported [1] describe the bfloat16 vdot & WMMA operations, and these are in fact implemented and working in various software such as composable kernels and rocBLAS, which I have used (and can guarantee they are not simply being run as float). I've also used these in the AMD fork of llm.c [3] Outside of gfx11, I have also used bfloat16 in CDNA2 & 3 devices, and they are working and being supported. Regarding cublasLt, what is your plan for support there? Pass everything through to hipblasLt (hipify style) or something else? Cheers, -A [1] https://llvm.org/docs/AMDGPU/AMDGPUAsmGFX11.html https://llvm.org/docs/AMDGPU/AMDGPUAsmGFX11.html [2] https://www.amd.com/content/dam/amd/en/documents/radeon-tech-docs/instruction-set-architectures/rdna3-shader-instruction-set-architecture-feb-2023_0.pdf https://www.amd.com/content/dam/amd/en/documents/radeon-tech... [3] http://github.com/anthonix/llm.c http://github.com/anthonix/llm.c
- ashvardanian 2y agoIt’s great that there is a page about current limitations [1], but I am afraid that what most people describe as “CUDA” is a small subset of the real CUDA functionality. Would be great to have a comparison table for advanced features like warp shuffles, atomics, DPX, TMA, MMA, etc. Ideally a table, mapping every PTX instruction to a direct RDNA counterpart or a list of instructions used to emulate it. [1]: https://docs.scale-lang.com/manual/differences/ https://docs.scale-lang.com/manual/differences/
- ckitching 2y agoYou're right that most people only use a small subset of cuda: we prioritied support for features based on what was needed for various open-source projects, as a way to try to capture the most common things first. A complete API comparison table is coming soon, I belive. :D In a nutshell: - DPX: Yes. - Shuffles: Yes. Including the PTX versions, with all their weird/wacky/insane arguments. - Atomics: yes, except the 128-bit atomics nvidia added very recently. - MMA: in development, though of course we can't fix the fact that nvidia's hardware in this area is just better than AMD's, so don't expect performance to be as good in all cases. - TMA: On the same branch as MMA, though it'll just be using AMD's async copy instructions. > mapping every PTX instruction to a direct RDNA counterpart or a list of instructions used to emulate it. We plan to publish a compatibility table of which instructons are supported, but a list of the instructions used to produce each PTX instruction is not in general meaningful. The inline PTX handler works by converting the PTX block to LLVM IR at the start of compilation (at the same time the rest of your code gets turned into IR), so it then "compiles forward" with the rest of the program. As a result, the actual instructions chosen vary on a csae-by-case basis due to the whims of the optimiser. This design in principle produces better performance than a hypothetical solution that turned PTX asm into AMD asm, because it conveniently eliminates the optimisation barrier an asm block typically represents. Care, of course, is taken to handle the wacky memory consistency concerns that this implies! We're documenting which ones are expected to perform worse than on NVIDIA, though!
- ashvardanian 2y agoHave you seen anyone productively using TMA on Nvidia or async instructions on AMD? I’m currently looking at a 60% throughput degradation for 2D inputs on H100: https://github.com/ashvardanian/scaling-democracy/blob/a8092613fac1ae107e6a956c5c41ad4994a51735/scaling_democracy.cu#L266 https://github.com/ashvardanian/scaling-democracy/blob/a8092...
- qwerty456127 2y ago> gfx1030, gfx1100, gfx1010, gfx1101, gfx900... How do I find out which do I have?
- ckitching 2y agoLike this: https://docs.scale-lang.com/manual/how-to-use/#identifying-gpu-target https://docs.scale-lang.com/manual/how-to-use/#identifying-g...
- systemBuilder 2y agogfx1101 : https://www.techpowerup.com/gpu-specs/amd-navi-32.g1000 https://www.techpowerup.com/gpu-specs/amd-navi-32.g1000 gfx1100 : https://www.techpowerup.com/gpu-specs/amd-navi-31.g998 https://www.techpowerup.com/gpu-specs/amd-navi-31.g998 gfx1030 : https://www.techpowerup.com/gpu-specs/amd-navi-21.g923 https://www.techpowerup.com/gpu-specs/amd-navi-21.g923 gfx1010 : https://www.techpowerup.com/gpu-specs/amd-navi-10.g861 https://www.techpowerup.com/gpu-specs/amd-navi-10.g861 gfx900 : https://www.techpowerup.com/gpu-specs/amd-vega-10.g800 https://www.techpowerup.com/gpu-specs/amd-vega-10.g800
- qwerty456127 2y agoThank you, I found out I have Vega 6 which apparently is gfx902.
- galaxyLogic 2y agoCompanies selling CUDA software should no doubt adopt this tool
- yieldcrv 2y agothe real question here is whether anybody has gotten cheap, easily available AMD GPUs to run their AI workloads, and if we can predict more people will do so
- JonChesterfield 2y agoMicrosoft have their production models running on amdgpu. I doubt it was easy but it's pretty compelling as an existence proof
- anthonix1 2y agoI ported Karparthy's llm.c repo to AMD devices [1], and have trained GPT2 from scratch with 10B tokens of fineweb-edu on a 4x 7900XTX machine in just a few hours (about $2 worth of electricity) [2]. I've also trained the larger GPT2-XL model from scratch on bigger CDNA machines. Works fine. [1] https://github.com/anthonix/llm.c https://github.com/anthonix/llm.c [2] https://x.com/zealandic1 https://x.com/zealandic1
- EGreg 2y agoBut the question is, can it also run SHUDA and WUDA?
- nabogh 2y agoI've written a bit of CUDA before. If I want to go pretty bare-bones, what's the equivalent setup for writing code for my AMD card?
- JonChesterfield 2y agoHIP works very similarly. Install rocm from your Linux distribution or from amd's repo, or build it from github.com/rocm. Has the nice feature of being pure userspace if you use the driver version that's already in your kernel. How turn-key / happy an experience that is depends on how closely your system correlates with one of the documented/tested distro versions and what GPU you have. If it's one that doesn't have binary versions of rocblas etc in the binary blob, either build rocm from source or don't bother with rocblas.
- spfd 2y agoVery impressive! But I can't help but think if something like this can be done to this extend, I wonder what went wrong/why it's a struggle for OpenCL to unify the two fragmentized communities. While this is very practical and has a significant impact for people who develop GPGPU/AI applications, for the heterogeneous computing community as a whole, relying on/promoting a proprietary interface/API/language to become THE interface to work with different GPUs sounds like bad news. Can someone educate me on why OpenCL seems to be out of scene in the comments/any of the recent discussions related to this topic?
- vedranm 2y agoIf you are going the "open standard" route, SYCL is much more modern than OpenCL and also nicer to work with.
- JonChesterfield 2y agoOpencl gives you the subset of capability that a lot of different companies were confident they could implement. That subset turns out to be intensely annoying to program in - it's just the compiler saying no over and over again. Or you can compile as freestanding c++ with clang extensions and it works much like a CPU does. Or you can compile as cuda or openmp and most stuff you write actually turns into code, not a semantic error. Currently cuda holds lead position but it should lose that place because it's horrible to work in (and to a lesser extent because more than one company knows how to make a GPU). Openmp is an interesting alternative - need to be a little careful to get fast code out but lots of things work somewhat intuitively. Personally, I think raw C++ is going to win out and the many heterogeneous languages will ultimately be dropped as basically a bad idea. But time will tell. Opencl looks very DoA.
- varelse 2y ago[dead]
- mschuetz 2y agoOpenCL isn't nice to use and lacks tons of quality of life features. I wouldn't use it, even if it was double as fast as CUDA.
- localfirst 2y ago> SCALE does not require the CUDA program or its build system to be modified. how big of a deal is this?
- JonChesterfield 2y agoPeople can be wildly hostile to changing their programs. The people who wrote it aren't here any more, the program was validated as-is, changing it tends to stop the magic thing working and so forth. That changing the compiler is strongly equivalent to changing the source doesn't necessarily influence this pattern of thinking. Customer requests to keep the performance gains from a new compiler but not change the UB they were relying on with the old are definitely a thing.
- rjurney 2y agoIf it's efficient, this is very good for competition.
- ekelsen 2y agoA major component of many CUDA programs these days involves NCCL and high bandwidth intra-node communication. Does NCCL just work? If not, what would be involved in getting it to work?
- pjmlp 2y agoThis targets CUDA C++, not CUDA the NVIDIA infrastructure for C, C++, Fortran, and anything else targeting PTX.
- ckitching 2y agoThe CUDA C APIs are supported as much in C as in C++ using SCALE! Cuda-fortran is not currently supported by scale since we haven't seen much use of it "in the wild" to push it up our priority list.
- anon291 2y agoIt doesn't matter though. NVIDIA distributes tons of libraries built atop CUDA that you cannot distribute or use on AMD chips legally. Cutlass, CuBLAS, NCCL, etc.
- tama_sala 2y agoCorrect, which one of the main moats Nvidia has when it comes to training
- ckitching 2y agoSCALE doesn't use cuBlas and friends. For those APIs, it uses either its own implementations of the functions, or delegates to an existing AMD library (such as rocblas). It wouldn't even be technically possible for SCALE to distribute and use cuBlas, since the source code is not available. I suppose maybe you could do distribute cuBlas and run it through ZLUDA, but that would likely become legally troublesome.
- anon291 2y ago> SCALE doesn't use cuBlas and friends. For those APIs, it uses either its own implementations of the functions, or delegates to an existing AMD library (such as rocblas). And this is the problem. I guarantee you NVIDIA has more engineers working on cuBLAS et al than AMD does. The NVIDIA moat is not CUDA the language or CUDA the library. It's CUDA the ecosystem. That means things like all the high performance libraries; all the high performance libraries with clustering support (does AMD even have a clustering solution like NVLink -- everyone forgets that NVIDIA also does high speed networking); all the high perf appliances (everyone also forgets that NVIDIA sells entire systems, not GPUS); all the high perf servers (Triton inference server, etc). We can go on. I commend the project volunteers for what they've done, but I would recommend getting VC money and competing directly with NVIDIA.
- uptownfunk 2y agoVery clearly the business motive make sense, go after nvidia gpu monopoly. Can someone help a lay person understand the pitfalls here that prevent this from being an intelligent venture?
- JonChesterfield 2y agoIt's technically non-trivial and deeply irritating to implement in places as people expect bugward compatibility with cuda. Also nvidia might savage you with lawyers for threatening their revenue stream. Big companies can kill small ones by strangling them in the courts then paying the fine when they lose a decade later.
- einpoklum 2y agoAt my workplace, we were reluctant in making the choice between writing OpenCL and being AMD-compliant, but missing out on CUDA features and tooling; and writing CUDA and being vendor-locked. Our jerry-rigged solution for now is writing kernels that are the same source for both OpenCL and CUDA, with a few macros doing a bit of adaptation (e.g. the syntax for constructing a struct). This requires no special library or complicated runtime work - but it does have the downside of forcing our code to be C'ish rather than C++'ish, which is quite annoying if you want to write anything that's templated. Note that all of this regards device-side, not host-side, code. For the host-side, I would like, at some point, to take the modern-C++ CUDA API wrappers (https://github.com/eyalroz/cuda-api-wrappers/ https://github.com/eyalroz/cuda-api-wrappers/) and derive from them something which supports CUDA, OpenCL and maybe HIP/ROCm. Unfortunately, I don't have the free time to do this on my own, so if anyone is interested in collaborating on something like that, please drop me a line. ----- You can find the OpenCL-that-is-also-CUDA mechanism at: https://github.com/eyalroz/gpu-kernel-runner/blob/main/kernels/include/port_from_cuda.cl.h https://github.com/eyalroz/gpu-kernel-runner/blob/main/kerne... and https://github.com/eyalroz/gpu-kernel-runner/blob/main/kernels/include/port_from_opencl.cuh https://github.com/eyalroz/gpu-kernel-runner/blob/main/kerne... (the files are provided alongside a tool for testing, profiling and debugging individual kernels outside of their respective applications.)
- JonChesterfield 2y agoFreestanding c++ with compiler intrinsics is a nicer alternative. You can do things like take the address of a function. Use an interface over memory allocation/queue launch with implementations in cuda, hsa, opencl whatever. All the rest of the GPU side stuff is syntax sugar/salt over slightly weird semantics, totally possible to opt out of all of that.
- stuaxo 2y agoWhat's the licensing, will I be able run this as a hobbyist for free software?
- tallmed 2y agoI wonder if this thing has anything common with zluda, its permissively licensed after all.
- EGreg 2y agoDoes it translate to OpenCL? This sounds like DirectX vs OpenGL debate when I was younger lol
- lukan 2y agoOk, so I just stumbled on the problem, that I tried out openwhisper (from OpenAI), but on my CPU, because of no CUDA and workarounds seem hacky. So the headline sounds good! But can this help me directly? Or would OpenAI have to use this tool for me to benefit? It is not immediately clear to me (but I am a beginner in this space).
- omneity 2y agoWondering if there's an ongoing effort to do the same with MPS/Metal as a backend. If anything given how many developers are on macs I think it could get immense traction.
- seanp2k2 2y agoThis is the way.
- qeternity 2y agoThe future is inference. Many inference stacks already support AMD although the kernels are less optimized. This will of course change over time, but if AMD can crack the inference demand, it will put NVDA under huge pressure.
- amelius 2y agoCan anyone explain why libcudnn is taking several gigabytes of my harddrive?
- gssa 2y ago[flagged]