6 ms·
The hardware is the easy part of accelerating NN training. Nvidia's software and infrastructure is so well designed and established that no competitor can threa
by WithinReason 1y ago
The hardware is the easy part of accelerating NN training. Nvidia's software and infrastructure is so well designed and established that no competitor can threaten them even if they give away the hardware for free.
- saagarjha 1y agoI don't know about well designed but it's definitely established.
- CoastalCoder 1y agoCould you elaborate? I've only done a little work on CUDA, but I was pretty impressed with it and with their NSys tools. I'm curious what you wish was different.
- saagarjha 1y agoI actually really hate CUDA's programming model and feel like it's too low-level to actually get any productive work done. I don't really blame Nvidia because they basically invented the programmable GPU and it wouldn't be fair to have them also come up with the perfect programming model right out of the gate but at this point it's pretty clear that having independent threads work on their own programs makes no sense. High performance code requires scheduling across multiple threads in a way that is completely different if you are coming from CPUs. Of course, one might mention that GPUs are nothing like CPUs–but the programming model works super hard to try to hide this. So it's not really well designed in my book. I actually quite like the compilers that people are designing these days to write block-level code, because I feel like it better represents the work people want to do and then you pick which way you want it lowered. As for Nsight (Systems), it is…ok, I guess? It's fine for games and stuff I guess but for HPC or AI it doesn't really surface the information that you would want. People who are running their GPUs really hard know they have kernels running all the time and what the performance characteristics of them are. Nsight Compute is the thing that tells you that but it's kind of a mediocre profiler (some of this may be limitations of hardware performance counters) and to use it effectively you basically have to read a bunch of blog posts by people instead of official documentation. Despite not having used it much, my impression was that Nvidia's "moat" was that they have good networking libraries, that they are pretty good (relatively) and making sure all their tools work, and they have had consistent investment on this for a decade.
- electroglyph 1y agoi mean, it could be worse... it could be Vulkan
- jandrewrogers 1y agoGPUs are a type of barrel processor, which are optimized for workloads without cache locality. As a fundamental principle, they replace the CPU cache with latency hiding behavior. Consequently, you can't use algorithms and data structures designed for CPUs, since most of those assume the existence of a CPU cache. Some things are very cheap on a barrel processor that are very expensive on a CPU and vice versa, which changes the way you think about optimization. The wide vectors on GPUs are somewhat irrelevant. Scalar barrel processors exist and have the same issues. A scalar barrel processor feels deceptively CPU-like and will happily compile and run normal CPU code. The performance will nonetheless be poor unless the C++ code is designed to be a good fit for the nature of a barrel processor, code which will look weird and non-idiomatic to someone who has only written code for CPUs. There is no way to hide that a barrel processor is not a CPU even though they superficially have a lot of CPU-like properties. A barrel processor is extremely efficient once you learn to write code for them and exceptionally well-suited to HPC since they are not latency-sensitive. However, most people never learn how to write proper code for barrel processors. Ironically, barrel processor style code architecture is easy to translate into highly optimized CPU code, just not the reverse.
- glitchc 1y agoI wanted to upvote you originally, but I'm afraid this is not correct. A GPU is not a barrel processor. In a barrel processor a single context is switched between multiple threads after each instruction. A barrel processor design has a singular instruction pipeline and a singular cache across all threads. In a GPU, due to the independence of the execution units, those threads will execute those instructions concurrently on all cores, as long as a program-based instruction dependency between threads is not introduced. It's true parallelism. Furthermore, each execution unit embeds its own instruction scheduler, it's own pipeline and its own L1 cache (see [1] for NVidia's architecture). [1] https://docs.nvidia.com/deeplearning/performance/dl-performance-gpu-background/index.html#gpu-arch https://docs.nvidia.com/deeplearning/performance/dl-performa...
- WithinReason 1y agoWho has better software than Nvidia for NN training? Meaning the least amount of friction getting a new network to train.
- saagarjha 1y agoJust because their tools are the best doesn't mean they are designed well.
- jacquesm 1y agoI've used DSPs, custom boards with compute hardware (FPGA image processing), and various kinds of GPUs. I would have a very hard time trying to point to ways in which the NVIDIA toolkit could be compared to what's out there and not come away with a massive sense of relief. For the most part 'it just works', the models are generic enough that you can actually get pretty close to the TDP on your own workloads with custom software and yet specific enough that you'll find stuff that makes your work easier most of the time. I really can't complain, now, FPGAs, however... And if there ever is a company that comes out and improves substantially on this I'll be happy for sure but if you asked me off the bat what they should improve I honestly wouldn't know, especially not taking into account that this was an incremental effort over ~2 decades and that originated in an industry that has nothing to do with the main use case today and some detours into unrelated industries besides (crypto, for instance). From fluid dynamics, FEA, crypto, gaming, genetics, AI and many others with a single generic architecture and delivering very good performance is no mean feat. I'd love to hear in what way you would improve on their toolset.
- programjames 1y agoNot the guy you replied to, but here are some improvements that feel obvious: 1. Memory indexing. It's a pain to avoid banking conflicts, and implement cooperative loading on transposed matrices. To improve this, (1) pop up a warning when banking conflicts are detected, (2) make cooperative loading solved by the compiler. It wouldn't be too hard to have a second form of indexing memory_{idx} that the compiler solves a linear programming problem for to maximize throughput (do you spend more thread cycles cooperative loading, or are banking conflicts fine because you have other things to work on?) 2. Why is there no warning when shared memory is unspecified? It isn't hard to check if you're accessing an index that might not have been assigned a value. The compiler should pop out a warning and assign it to 0.0, or maybe even just throw an error. 3. Timing - doesn't exist. Pretty much the gold standard is to run your kernel 10_000 times in a loop and subtract the time from before and after the loop. This isn't terribly important, I'm just getting flashbacks to before I learned `timeit` was a thing in Python.
- KeplerBoy 1y agoIt's not all about NNs and AI. Take a look at the Top500, a lot of people are doing classical HPC work on Nvidia GPUs, which are increasingly not designed for this. Unfortunately the HPC market is just a lot smaller than the AI bubble.
- rwmj 1y agoIf the hardware isn't available at all, we'll never find out if the software moat could be overcome.
- DrNosferatu 1y ago> if they give away the hardware for free. Seriously doubt that: free hardware (or 10s of bucks) would galvanize the community and achieve huge support - look at the Raspberry Pi project original prices and the consequences.
- deleted 1y ago[deleted]
- DrNosferatu 1y agoIn fact, if any such thing would happen, I would wager Nvidia stock would tank massively. Say, release has extensions to a RISC-V design.
- DrNosferatu 1y ago*as :D
- londons_explore 1y agoThe math of NN training isn't complex at all. Designing the software stack to make a new pytorch backend is very doable with the budgets these AI companies have. I suspect that whenever you look like you're making good progress on this front, nvidia gives you a lot of chips for free on condition you shelve the effort though! The latest example being Tesla, who were designing their own hardware and software stack for NN training, then suspiciously got huge numbers of H100's ahead of other clients and cancelled the dojo effort.
- AlotOfReading 1y agoI doubt that's what happened. They had designs that were massively expensive to fab/package, had much worse performance than the latest Nvidia hardware, and still needed massive amounts of custom in-house development. To combat all of these issues, they were fighting with Nvidia (and losing) for access to leading edge nodes, which kept going up in price. Their personnel costs kept rising as the company became more politicized, people left to join other companies (e.g. densityai), and they became embroiled in the salary wars to replace them. My suspicion is that Musk told them to just buy Nvidia instead of waiting around for years of slow iteration to get something competitive. The custom silicon I was involved with experienced similar issues. It was too expensive and slow to try competing with Nvidia, and no one could stomach the costs to do so.
- deleted 1y ago[deleted]
- nromiun 1y agoI don't know why you are getting downvoted. This is 100% true. It's not like you can take any random data and train it into a NN. You have to transform the data, you have to write the low level GPU kernels which will actually run fast on that particular GPU, you also have to get the output and transform that as well. All of this is hard and very much impossible to create from scratch. If people use PyTorch on a Nvidia GPU they are running layers and layers of code written by those that know how to write fast kernels for GPUs. In some cases they use assembly as well. Nvidia stuck to one stack and wrote all their high level libraries on it, while their competitors switched from old APIs to new ones and never made anything close to CUDA.
- woooooo 1y agoBecause in the context of LLM transformers, you really just need matrix multiplication to be hyper-optimized, it's 90-99% (citation needed) of the FLOPs. Get some normalization and activation functions in and you're good to go. It's not a massive software ecosystem. CUDA and CUBLAS being capable of a bunch of other things is really cool, and would take a long time to catch up with, but getting the bare minimum to run LLMs on any platform with a bunch of GDDR7 channels and cores at a reasonable price would have people writing torch/ggml backends within weeks.
- nromiun 1y agoHave you tried to write a kernel for basic matrix multiplication? Because I have and I can assure you it is very hard to get 50% of maximum FLOPs, let alone 90%. It is nothing like CPUs where you write a * b in C and get 99% of the performance by the compiler. Here is an example of how hard it is: https://siboehm.com/articles/22/CUDA-MMM https://siboehm.com/articles/22/CUDA-MMM And this is just basic matrix mult. If you add activation functions it will slow down even more. There is nothing easy about GPU programming, if you care about performance. CUDA gives you all that optimization on a plate.
- woooooo 1y agoWell, CUDA gives you a whole programming language where you have to figure out the optimization for your particular card's cache size and bus width. I'm saying the API surface of what to offer for LLMs is pretty small. Yeah, optimizing it is hard but it's "one really smart person works for a few weeks" hard, and most of the tiling techniques are public. Speaking of which, thanks for that blog post, off to read it now.