6 ms·
CUDA Tile Open Sourced
- jauntywundrkind 10mo agoWill be interesting to see if Nvidia and other have any interest & energy getting this used by others, if there actually is an ecosystem forming around it. Google leading XLA & IREE, with awesome intermediate representations, used by lots of hardware platforms, and backing really excellent Jax & Pytorch implementations, having tools for layout & optinization folks can share: they really build an amazing community. There's still so much room for planning/scheduling, so much hardware we have yet to target. RISC-V has really interesting vector instructions, for example, and it seems like there's so much exploration / work to do to better leverage that. Nvidia has partners everywhere now. Nvlink is used by Intel, AWS Tritanium, others. Yesterday the Groq exclusive license that Nvidia paid to give to Groq?! Seeing how and when CUDA Tiles emerges: will be interesting. Moving from fabric partnerships, up up up the stack.
- turtletontine 10mo agoOn the RISC-V vector instructions, could you elaborate? Are the vector extensions substantially different from those in ARM or x86?
- adgjlsfhk1 10mo agoit's fairly similar to Arm's sve2, but very different from the x86 side in that the instructions are variable length rather than fixed
- Moosdijk 10mo ago> There's still so much room for planning/scheduling, so much hardware we have yet to target this is nicely illustrated by this recent article: https://news.ycombinator.com/item?id=46366998 https://news.ycombinator.com/item?id=46366998
- pjmlp 10mo agoFor NVidia it suffices this is a Python JIT allowing programming CUDA compute kernels directly in Python instead of C++, yet another way how Intel and AMD, alongside Khronos APIs, lag behind in great developer experiences for GPU compute programming. Ah, and Nsight debugging also supports Python CUDA Tiles debugging. https://developer.nvidia.com/blog/simplify-gpu-programming-with-nvidia-cuda-tile-in-python/ https://developer.nvidia.com/blog/simplify-gpu-programming-w...
- Q6T46nT668w6i3m 10mo agoSlang is a fantastic developer experience.
- saagarjha 10mo agoNsight does not have a debugger.
- dahart 10mo agoWhat do you mean? Are you unaware of Nsight VSE? https://developer.nvidia.com/nsight-visual-studio-edition https://developer.nvidia.com/nsight-visual-studio-edition
- saagarjha 10mo agoI was aware of their Visual Studio plugins but I did not know that they called their debugger support for Visual Studio “Nsight” as well.
- almostgotcaught 10mo ago> Google leading XLA & IREE IREE hasn't been at G for >2 years.
- nl 10mo ago> Groq exclusive license non-exclusive license actually.
- CamperBob2 10mo agoFun game: see how many clicks it takes you to learn what MLIR stands for. I lost count at five or six. Define your acronyms on first use, people.
- fragmede 10mo agoI did it in three. I selected it in your comment, and then had to hit "more" to get to the menu to ask Google about it, which brought me to https://www.google.com/search?q=MLIR https://www.google.com/search?q=MLIR which says: MLIR is an open-source compiler infrastructure project developed as a sub-project of the LLVM project. Hopefully Get better at computers and stop needing to be spoon-fed information, people!
- reactordev 10mo agoIn this day and age, asking questions about what something is is a minefield of “just ask AI” and “You should know this”. Let’s stop putting down people who ask questions and root out those that have shitty answers.
- ThrowawayTestr 10mo agoGoogle is nearly 30 years old
- pjmlp 10mo agoAnd we are not counting Yahoo, Altavista, Ask Jeeves, MSN,...
- fragmede 10mo agoI get why it feels frustrating when someone snaps "just google it." Nobody likes feeling dumb. That said, there’s a meaningful difference between asking a genuine question and demanding that every discussion be padded to accommodate readers who won’t even type four letters into a search bar. Expecting complete spoon-feeding in technical threads isn’t curiosity; it’s a refusal to engage. Learning requires participation.
- xmorse 10mo agoWriting this in Mojo would have been so much easier
- 3abiton 10mo agoIt's barely gaining adoption though. The lack of buzz is a chicken and egg issue for Mojo. I fiddled shortly with it (mainly to get it working some of my pythong scripts), and it was suprisingly easy. It'll shoot up one day for sure if Latner doesn't give up early on it.
- ronsor 10mo agoIsn't the compiler still closed source? I and many other ML devs have no interest in a closed-source compiler. We have enough proprietary things from NVIDIA.
- 0x696C6961 10mo agoYeah, the mojo pitch is so good, but I don't think anyone has an appetite for the potential fuckery that comes with a closed source platform.
- 3abiton 10mo agoYes, but Latner said multiple time it's closed until it matures (he apparently did this with llvm and swift too). So not unusal. His open source target is end of 2026. In all fairness, I have 0 doubts that he would deliver.
- boywitharupee 10mo agoshouldn't the title be "CUDA Tile IR Open Sourced"?
- OneDeuxTriSeiGo 10mo agoIt's more or less the same thing. CUDA TIle is the name of the IR, cuTile is the name of the high level DSLs.
- toolboxg1x0 10mo agoNVIDIA tensor core units, where the second column in kernel optimization is producing a test suite.
- opan 10mo ago>The CUDA Tile IR project is under the Apache License v2.0 with LLVM Exceptions
- deleted 10mo ago[deleted]
- fooblaster 10mo agoLet's see if developers sleepwalk into another trap to keep us locked into nvidia's hardware for the next decade.
- the__alchemist 10mo agoIMO it's not Nvidia's fault the competing APIs are high friction.
- flyingcoder 10mo agoAMD screwed up so badly.
- fooblaster 10mo agoThat is true, but that doesn't mean Nvidia is not engaging in engineering to intentionally kneecap competition. Triton and other languages like that are a huge threat and CUtile is a means to combat that threat and prevent a hardware abstraction layer.
- positron26 10mo agoHundreds of thousands of developers with access to a global communication network were not stopped by AMD. Why act like dependents or wait for some bright star of consensus unless the intent is really about getting the work for free? We don't have to wait for singular companies or foundations to fix ecosystem problems. Only the means of coordination are needed. https://prizeforge.com https://prizeforge.com isn't there yet, but it is already capable of bootstrapping its own development. Matching funds, joining the team, or contributing on MuTate will all make the ball pick up speed faster.
- nemothekid 10mo ago>We don't have to wait for singular companies or foundations to fix ecosystem problems. Geohot has been working on this for about a year, and every roadblock he's encountered he has had to damn near pester Lisa Su about getting drivers fixed. If you want the CUDA replacement that would work on AMD, you need to wait on AMD. If there is a bug in the AMD microcode, you are effectively "stopped by AMD".
- gaogao 10mo agoThe compiler for CUDA Tile being Blackwell only is a baffling decision. I wanted to try it out, but it's only really easy to grab H100s quickly right now. I guess maybe I'll try it out on my 5070 Ti after traveling, but am more likely to stick to an IR that targets multiple platforms, since they couldn't be bothered.
- robobsolete 10mo agoI was keen to try it too, but oh well
- 0-_-0 10mo agoThis is basically the nvidia equivalent of cooperative_matrix_2 in Vulkan which is vendor agnostic and should get much more hype that it's getting.
- pyuser583 9mo agoI’m glad CUDA and “open source” are in the same sentence again. We’d all prefer cross platform programming, but if you’re going to do platform specific, I prefer open source to closed source. Thank you NVIDIA!