5 ms·
NVIDIA does not want CUDA development (e.g. flash attention) to move to Triton because Triton also supports AMD and if ecosystem moves from pure CUDA to Triton,
by lukax 1y ago
NVIDIA does not want CUDA development (e.g. flash attention) to move to Triton because Triton also supports AMD and if ecosystem moves from pure CUDA to Triton, that's bad for NVIDIA's lock-in. That's why there is so much focus on CUDA Python (lower level) and Tilus (higher level, more similar to Triton).
- archerx 1y agoWhat’s bad for Nvidia is good for everyone else. The cuda lock-in needs to die.
- pjmlp 1y agoIt is on AMD and Intel to deliver.
- ants_everywhere 1y agoOr one of the cloud providers who doesn't want to pay lock-in prices when they'd rather pay commodity prices
- topspin 1y agoHow feasible is this for a cloud operation(s)? I imagine this work requires close collaboration with the architects and proprietary knowledge about the design.
- ants_everywhere 1y agoit seems feasible, it's more a matter of how much of a priority it is. I follow Google most closely. They design and manufacture their own accelerators. AWS I know manufactures its own CPUs, but I don't know if they're working on or already have an AI accelerator. Several of the big players are working on OpenXLA, which is designed to abstract and commoditize the GPU layer: https://openxla.org/xla https://openxla.org/xla OpenXLA mentions: > Alibaba, Amazon Web Services, AMD, Apple, Arm, Google, Intel, Meta, and NVIDIA
- mdaniel 1y ago> AWS I know manufactures its own CPUs, but I don't know if they're working on or already have an AI accelerator I believe those are the Inferentia: https://aws.amazon.com/ai/machine-learning/inferentia/ https://aws.amazon.com/ai/machine-learning/inferentia/ > AWS Inferentia chips are designed by AWS to deliver high performance at the lowest cost in Amazon EC2 for your deep learning (DL) and generative AI inference applications but I don't know this second if they're supported by the major frameworks, or what I also didn't recall about https://aws.amazon.com/ai/machine-learning/trainium/ https://aws.amazon.com/ai/machine-learning/trainium/ until I was looking up that page, so it seems they're trying to have a competitor to the TPUs just naming them dumb, because AWS > AWS Trainium chips are a family of AI chips purpose built by AWS for AI training and inference to deliver high performance while reducing costs.
- ants_everywhere 1y agothanks this is useful! > have a competitor to the TPUs just naming them dumb, because AWS I kind of like "trainium" although "inferentia" I could take or leave. At least it's nice that the names tell you the intended use case.
- pjmlp 1y agoWith what software though?
- Twirrim 1y agoNot sure cloud providers will care, all the costs get passed onto the customers. There's already far more demand for GPUs than can be met by the supply chain, too. If they were sitting on excess stock, or struggling to sell, sure.
- coredog64 1y agoThe cloud providers all have their own Nvidia alternatives. Having worked with more than one, I would rate them not much better than AMD when it comes to software.
- nabla9 1y agoAnd they continue to fumble with it. AMD has had time to catch up--a decade in fact. They simply don’t understand: robust software support requires a significant investment from their side. Simply providing small amounts of funding for academic research doesn’t suffice. Meanwhile Nvidia keeps building more and more libraries..
- DSingularity 1y agoIt’s not AMD it’s their board. Unless the board approves billions of $ in stock rewards to motivate good engineers nobody is going to join. It’s not rocket science. They can identify many key personel in Nvidia and make them offers which would be significantly better for them. Cycle 3 years and repeat. Two or three cycles and you will have replicated the most important parts.
- teeklp 1y agoIt wasn't me that missed the deadline, it was my brain.
- gary_0 1y agoIt's my guess that they don't want to. A decade ago when AMD's CPUs had trouble competing at the high end, they ceded that market segment almost entirely, and now they're doing the same with nVidia. And Intel is basically a dead company. Neither of them are going to risk the capital and internal shake-up necessary to actually compete with nVidia. And anyways, there's only so much TSMC capacity for high-end chips, and Apple and nVidia have already spent infinity dollars reserving most of it.
- hgehjddfy 1y agoUmmm are you forgetting epyc and ryzen are now unmatched? If AMD wants to they can compete..
- gary_0 1y agoUnmatched by Intel, who have been failing for a decade, so there was no competition? They mostly won by default. And Apple's chips are giving them a run for their money. If Apple sold plain CPUs that weren't locked to their software (they never will, but hypothetically) then AMD would let themselves slide into 2nd place again. That really makes three companies that are happy to concede to nVidia, because Apple could definitely challenge nVidia if they wanted to. Note: I'm not saying that AMD sucks, just that their corporate culture prevents them from being very ambitious.
- AuthAuth 1y agoApples chips dont even come close. Their benchmarks compete in specific tasks and then measure by metrics like preformance per watt. These are benchmarks AMDs cpus arent optimizing for and yet they're still close. Once you remove the power consumption out of the tests and broading the tests AMD cpu's come out ahead. Apple had something impressive with M1 then within a year the other mobile cpu manufactures came out with something on par. A year after that and they had surpassed apple. Apples closest cpu competition is Qualcomm and they dont win that.
- dismalaf 1y agoWhat's crazy to me is that ROCm and SYCL are open-source, but somehow more difficult to install and support less hardware for their respective brands than CUDA does for Nvidia...
- torginus 1y agoAfter looking up Triton this seems quite different - Triton is a high-level CUDA competitor with Python-like syntax, and this seems to be a library aimed at generating GPU assembly for micro-optimizing kernels.
- YetAnotherNick 1y agoLooks pretty similar to me[1] [1]: https://nvidia.github.io/tilus/getting-started/tutorials/matmul/matmul_v1.html https://nvidia.github.io/tilus/getting-started/tutorials/mat...
- mdaniel 1y agoMy experience searching for "nvidia triton" coughed up oppressive number of results for the similarly named inference server but I think the Triton discussed here is https://github.com/triton-lang/triton https://github.com/triton-lang/triton (although the commits that I spot checked were all from @openai.com emails, not nvidia)