11 ms·
Triton: Open-Source GPU Programming for Neural Networks
- notthedroids 5y agoDoes Triton support automatic differentiation? I don't see that feature in a quick poke through the docs. If it does compile to LLVM, I suppose it can use Enzyme https://enzyme.mit.edu/ https://enzyme.mit.edu/
- bmh 5y agoThis sits BELOW the automatic differentiation layer.
- riyadparvez 5y agoI am confused. Is it another competitor of Tensorflow, JAX, and Pytorch? Or something else?
- peytoncasper 5y agoI believe this is more of an optimization layer to be utilized by libraries like Tensorflow and JAX. More of a simplification of the interaction with traditional CUDA instructions. I imagine these libraries and possibly some users would implement libraries on top of this language and reap some of the optimization benefit without having to maintain low-level CUDA specific code.
- blueblisters 5y agoSo is this similar to XLA?
- jpf0 5y agoXLA is domain-specific compiler for linear algebra. Triton generates and compiles an intermediate representation for tiled computation. This IR allows more general functions and also claims higher performance. obligatory reference to the family of work: https://github.com/merrymercy/awesome-tensor-compilers https://github.com/merrymercy/awesome-tensor-compilers
- anon_tor_12345 5y agoWithout reading the paper, I think you have it a little backwards - the IR doesn't itself allow for more general functions. More general functions are possible (in theory) because the frontend (this Triton language) is decoupled from the backend (CUDA) through the IR as an interface. In this way the Triton IR is no less domain specific than XLA (because both are IRs that represent sequences of operators that run on GPU (or TPU or whatever). I guess in theory Triton could be eschewing all of eg cuDNN but most likely it's not as NVIDIA's closed source kernels perform best on their closed source hardware. Edit: should've read the post before commenting. Looks like they are in fact using LLVM's PTX backend (ie generating cuda kernels from scratch). Kudos to them
- jacoblambda 5y agoIm curious how it would compare to Halide Lang. They both seem to be targetting the same problem.
- sanxiyn 5y agohttps://triton-lang.org/programming-guide/chapter-2/related-work.html https://triton-lang.org/programming-guide/chapter-2/related-... has comparison with Halide, which is categorized as "Scheduling Languages" there.
- N1H1L 5y agoFrom what I got from reading the docs and the blog post, was that this is a competitor to torch.jit and numba.cuda.jit - write faster Pythonic GPGPU kernels without sacrificing speed.
- croes 5y agoToo bad it's CUDA Sooner or later this will become a problem because you are depending on the benevolence of a single manufacturer.
- yumraj 5y agoSure, but if this abstraction layer becomes popular, then it becomes much easier to support other GPUs without requiring client libraries to change, which is a much harder problem.
- dragontamer 5y agoA big reason why CUDA is popular with compilers is that the PTX assembly-ish language is well documented and reasonable. Compilers generate PTX, then the rest of the CUDA infrastructure turns PTX into Turing machine code, or Ampere machine code, or Pascal machine code. In theory, SPIR-V should do the same job, but its just not as usable right now. In the meantime, getting it to work on PTX is easier, and then there's probably hope (in the far future) to move to SPIR-V if that ever actually takes off. I'm not a developer on Triton, but that'd be my expectation.
- pjmlp 5y agoPTX and the immense tooling around CUDA. Khronos mindset of it must be C like and parterns will take care of the ecosystem is what doomed OpenCL. All of their API design endevours are "design by committe" at its best. No wonder that SYSCL is now backend agnostic.
- shubuZ 5y agoWhich other hardware vendor provides the level of performance that Nvidia's GPU provide? Wasnt the benevolence on single (or couple) manufacturer(s) true in 90s, 2020s?
- croes 5y agoIt's not about performance but open standards. Remember Oracle vs. Google, at some time in the future NVidia could decide to get money out of CUDA.
- boulos 5y agoFolks might find the author’s research paper [1] while at Harvard more informative. This is a great high-level description, but if you want more detail, I recommend the paper. [1] https://dl.acm.org/doi/abs/10.1145/3315508.3329973 https://dl.acm.org/doi/abs/10.1145/3315508.3329973
- lsb 5y agoThat's http://www.eecs.harvard.edu/~htk/publication/2019-mapl-tillet-kung-cox.pdf http://www.eecs.harvard.edu/~htk/publication/2019-mapl-tille... for those of us outside the paywall
- queuebert 5y agoAlways wait for the second link.
- boulos 5y agoHuh, I thought the ACM DL was open "now" during the pandemic. Apologies!
- mjn 5y agoThat was unfortunately short-lived. ACM announced on March 30, 2020 that they would open their DL for 90 days due to the pandemic [1]. I don't believe there was an extension, so it expired on June 30, 2020. [1] https://www.acm.org/articles/bulletins/2020/march/dl-access-during-covid-19 https://www.acm.org/articles/bulletins/2020/march/dl-access-...
- mshachkov 5y agoNote, that there is a new (substantially accelerated) "autoscheduler" implementation in TVM (one of competitors in the article linked), details can be found in https://arxiv.org/abs/2006.06762 https://arxiv.org/abs/2006.06762
- shubuZ 5y agoI have found writing CUDA code is much simpler than writing correct multi-threaded AVX2/AVX-512 code.
- dragontamer 5y agoIf you need CPU-side SIMD, then try ispc: https://ispc.github.io/ https://ispc.github.io/ Its pretty much the OpenCL-model, except it compiles into AVX2 code / AVX512 code. Very similar to CUDA / OpenCL style programming. Its not single-source like CUDA, but it largely accomplishes the programming model IMO.
- gnufx 5y agoWhy not a standard? OpenMP is more than pretty much C(++) and Fortran, and has offload inspired by the needs of the Sierra supercomputer.
- dragontamer 5y agoOpenMP 4.5+ looks very promising to me (particularly the "simd" keyword associated with for-loops). But open-source implementations of OpenMP are somewhat lackluster... at least last time I checked it out. Maybe its time I revisit it. I've always thought the OpenMP spec was being written by highly competent programmers. They seem to "get" what is needed. But the question is if I can get my hands on any OpenMP implementation to actually play with. Not all of us can afford IBM's compiler suite!
- jpf0 5y agoLLVM has an openMP implementation
- dragontamer 5y agoThe task-based parallelism in LLVM leaves much to be desired however. Ideally, you'd want a more efficient implementation. But yeah, good enough to play with. But maybe not good enough to achieve high levels of performance. The SIMD stuff is probably simple enough to implement... maybe I should checkout how well LLVM works with OMP SIMD keywords.
- mvanaltvorst 5y agoAs a sidenote, these SVG graphs are absolutely beautiful. The flow diagram is even responsive on mobile!
- thebruce87m 5y agoUnfortunate name clash with NVIDIAs Triton Inference Server: https://developer.nvidia.com/nvidia-triton-inference-server https://developer.nvidia.com/nvidia-triton-inference-server
- deleted 5y ago[deleted]
- polynomial 5y agoMy first thought exactly. This will cause nothing but confusion and Triton (the inference server) is well integrated into the space. So it's especially weird to see it coming from OpenAI, and not a more random startup. It honestly makes no sense they would deliberately do this, unless there is some secret cult of Triton that is going on in the Bay Area world of AI/ML.
- 6gvONxR4sf7o 5y agoThe author commented on reddit about that (https://www.reddit.com/r/MachineLearning/comments/otdpkx/n_introducing_triton_opensource_gpu_programming/h6uo9hr/ https://www.reddit.com/r/MachineLearning/comments/otdpkx/n_i...) > PS: The name Triton was coined in mid-2019 when I released my PhD paper on the subject (http://www.eecs.harvard.edu/~htk/publication/2019-mapl-tillet-kung-cox.pdf http://www.eecs.harvard.edu/~htk/publication/2019-mapl-tille...). I chose not to rename the project when the Triton inference server came out a year later since it's the only thing that ties my helpful PhD advisors to the project.
- king_magic 5y agoThe author is unfortunately wrong. NVIDIA's Triton was referenced in marketing material as far back as 2018. https://developer.nvidia.com/blog/nvidia-serves-deep-learning-inference/ https://developer.nvidia.com/blog/nvidia-serves-deep-learnin...
- wyldfire 5y agoConfusingly similar: check Same industry: check Maybe the author should expect an incoming C&D. IIRC US is first-to-use, so NVIDIA would prevail?
- giacaglia 5y agoOpenAI keeps innovating. Amazing to see the speed of execution of the team
- ipsum2 5y agoThis guy developed Triton for his PhD thesis, and OpenAI hired him to continue working on it. Doesn't really seem fair to give all the innovation credit to OpenAI. See: https://www.reddit.com/r/MachineLearning/comments/otdpkx/n_introducing_triton_opensource_gpu_programming/h6uo9hr/ https://www.reddit.com/r/MachineLearning/comments/otdpkx/n_i...
- xmaayy 5y agoI wonder if this can be used for graphics programming. Shaders are notoriously hard to write correctly and this seems like it might provide an easier gateway than OpenGLSL
- ipsum2 5y agoTaichi is a similar project focused on graphics: https://taichi.graphics/ https://taichi.graphics/
- fzimmermann89 5y agoSo the code looks (apart from pointers) similiar to numba which feels much closer to numpy/pytorch high level code. Are there huge advantages in the triton model compared to numba that I don't see? Or is there a big performance gap? For me numba was always the easiest way to get some new idea running on cuda, and most of the time it was fast enough.. Did anybody find performance comparison between numba and triton?
- sanxiyn 5y agoUnlike Numba, Triton operators operate on blocks with explicit load and store of blocks. This is what enables analysis to automate coalescing, shared memory management, etc. I guess Triton is not for you if Numba is fast enough.
- shmerl 5y agoThis seems to be tied to CUDA? Why not build it on top of portable GPU programming base?
- einpoklum 5y agoThis sounds like basically hand-holding for Python programmers to write simple NN-operations. I'm sure it's convenient and useful, but it's still glorified glue code.
- blt 5y agoThis is kind of the opposite of glue code.
- john579 5y agoVery perceptive, dead on.
- bmh 5y agoNo, this is a DSL that allows people who normally write CUDA, to do so with less lines of code, and end up with a faster kernel. By embedding it inside Python you don't need to write your own lexer/parser.
- JDDunn9 5y ago> CPUs and AMD GPUs are not supported at the moment, but we welcome community contributions aimed at addressing this limitation. That's disappointing. My biggest frustration is that every ML library (Pytorch, Keras, etc) is tied to CUDA/Nvidia, so I take a huge performance hit when running them on my Mac.
- ganoushoreilly 5y agoI'm really surprised AMD hasn't pushed much harder into the space given their aggressive targeting of the data center with Epyc. I too use mostly Macs during the day and any ML projects are always relegated to Nvidia machines in my rack / amazon.
- make3 5y agonot pytorch anymore. there's support now
- voldacar 5y agoIf python had proper macros/easy AST processing you wouldn't need to make this a separate language. still cool
- bmh 5y agoA toy illustrative example, summing two arrays: CUDA c[i] = a[i] + b[i] i += 1 Triton c[i:i+16] = a[i:i+16] + b[i:i+16] i += 16 The 16 in this example is the "block size", and could be anything. But this notion of expressing computation over blocks of dense data seems to be the big difference from other approaches. A very exciting result of the incredible performance that Triton achieves, is the ability to fuse NN operations such as Matrix Multiply + LeakyReLU + Batch Norm. Previously, you needed to rely on cuBLAS for fast hand-written Matrix Multiply kernels, and then your LeakyReLU would need to read that result out of memory, and then your Batch Norm would read the LeakyReLU out of memory again. The ability to write very fast kernels, and especially being able to fuse them together, to avoid unnecessary memory round-trips is a big deal!
- volta83 5y ago> Previously, you needed to rely on cuBLAS for fast hand-written Matrix Multiply kernels, and then your LeakyReLU would need to read that result out of memory, You could do that, but you can also just tell cuBLAS to fuse ReLU, by just passing the "CUBLASLT_EPILOGUE_RELU" option (among others), see the manual: https://docs.nvidia.com/cuda/cublas/index.html#cublasLtEpilogue_t https://docs.nvidia.com/cuda/cublas/index.html#cublasLtEpilo... This has been possible for years. It's the kind of 1 line change that makes a big difference.