24 ms·
How to Think About GPUs
- porridgeraisin 1y agoA short addition that pre-volta nvidia GPUs were SIMD like TPUs are, and not SIMT which post-volta nvidia GPUs are.
- camel-cdr 1y agoSIMT is just a programming model for SIMD. Modern GPUs still are just SIMD with good predication support at ISA level.
- porridgeraisin 1y agoI was referring to this portion of TFA > CUDA cores are much more flexible than a TPU’s VPU: GPU CUDA cores use what is called a SIMT (Single Instruction Multiple Threads) programming model, compared to the TPU’s SIMD (Single Instruction Multiple Data) model.
- adrian_b 1y agoThis flexibility of CUDA is a software facility, which is independent of the hardware implementation. For any SIMD processor one can write a compiler that translates a program written for the SIMT programming model into SIMD instructions. For example, for the Intel/AMD CPUs with SSE4/AVX/AVX-512 ISAs, there exists a compiler of this kind (ispc: https://github.com/ispc/ispc https://github.com/ispc/ispc).
- porridgeraisin 1y agoThanks, I will look into that. However, I'm still confused about the original statement. What I had thought was that pre-volta GPUs, each thread in a warp has to execute in lock-step. Post-volta, they can all execute different instructions. Obviously this is a surface level understanding. How do I reconcile this with what you wrote in the other comment and this one?
- achierius 1y agoThat's not true. SIMT notably allows for divergence and reconvergence, whereby single threads actually end up executing different work for a time, while in SIMD you have to always be in sync.
- camel-cdr 1y agoI'm not aware of any GPU that implements this. Even the interleaved execution introduced in Volta still can only execute one type of instruction at a time [1]. This feature wasn't meant to accelerate code, but to allow more composable programming models [2]. Going of the diagram, it looks equivilant to rapidly switching between predicates, not executing two different operations at once. if (theradIdx.x < 4) { A; B; } else { X; Y; } Z; The diagram shows how this executes in the following order: Volta: ->| ->X ->Y ->Z|-> ->|->A ->B ->Z |-> pre Volta: ->| ->X->Y|->Z ->|->A->B |->Z The SIMD equivilant of pre Volta is: vslt mask, vid, 4 vopA ..., mask vopB ..., mask vopX ..., ~mask vopY ..., ~mask vopZ ... The Volta model is: vslt mask, vid, 4 vopA ..., mask vopX ..., ~mask vopB ..., mask vopY ..., ~mask vopZ ... [1] https://chipsandcheese.com/i/138977322/shader-execution-reordering https://chipsandcheese.com/i/138977322/shader-execution-reor... [2] https://stackoverflow.com/questions/70987051/independent-thread-scheduling-since-volta https://stackoverflow.com/questions/70987051/independent-thr...
- namibj 1y agoIIUC volta brought the ability to run a tail call state machine with let's presume identically-expensive states and state count less than threads-per-warp, at an average goodput of more than one thread actually active. Before it would loose all parallelism as it couldn't handle different threads having truly different/separate control flow, emulating dumb-mode via predicated execution/lane-masking.
- adrian_b 1y ago"Divergence" is supported by any SIMD processor, but with various amounts of overhead depending on the architecture. "Divergence" means that every "divergent" SIMD instruction is executed at least twice, with different masks, so that it is actually executed only on a subset of the lanes (i.e. CUDA "threads"). SIMT is a programming model, not a hardware implementation. NVIDIA has never explained exactly how the execution of divergent threads has been improved since Volta, but it is certain that, like before, the CUDA "threads" are not threads in the traditional sense, i.e. the CUDA "threads" do not have independent program counters that can be active simultaneously. What seems to have been added since Volta is some mechanism for fast saving and restoring separate program counters for each CUDA "thread", in order to be able to handle data dependencies between distinct CUDA "threads" by activating the "threads" in the proper order, but those saved per-"thread" program counters cannot become active simultaneously if they have different values, so you cannot execute simultaneously instructions from different CUDA "threads", unless they perform the same operation, which is the same constraint that exists in any SIMD processor. Post-Volta, nothing has changed when there are no dependencies between the CUDA "threads" composing a CUDA "warp". What has changed is that now you can have dependencies between the "threads" of a "warp" and the program will produce correct results, while with older GPUs that was unlikely. However dependencies between the CUDA "threads" of a "warp" shall be avoided whenever possible, because they reduce the achievable performance.
- aanet 1y agoFantastic resource! Thanks for posting it here.
- deleted 1y ago[deleted]
- nickysielicki 1y agoThe calculation under “Quiz 2: GPU nodes“ is incorrect, to the best of my knowledge. There aren’t enough ports for each GPU and/or for each switch (less the crossbar connections) to fully realize the 450GB/s that’s theoretically possible, which is why 3.2TB/s of internode bandwidth is what’s offered on all of the major cloud providers and the reference systems. If it was 3.6TB/s, this would produce internode bottlenecks in any distributed ring workload. Shamelessly: I’m open to work if anyone is hiring.
- aschleck 1y agoIt's been a while since I thought about this but isn't the reason providers advertise only 3.2tbps because that's the limit of a single node's connection to the IB network? DGX is spec'ed to pair each H100 with a Connect-X 7 NIC and those cap out at 400gbps. 8 gpus * 400gbps / gpu = 3.2tbps. Quiz 2 is confusingly worded but is, iiuc, referring to intranode GPU connections rather than internode networking.
- charleshn 1y agoYes, 450GB/s is the per GPU bandwidth in the nvlink domain. 3.2Tbps is the per-host bandwidth in the scale out IB/Ethernet domain.
- jacobaustin123 1y agoI believe this is correct. For an H100, the 4 NVLink switches each have 64 ports supporting 25GB/s each, and each GPU uses a total of 18 ports. This gives us 450GB/s bandwidth within the node. But once you start trying to leave the node, you're limited by the per-node InfiniBand cabling, which only gives you 400GB/s out of the entire node (50GB / GPU).
- xtacy 1y agoIs it GBps (gigabytes per second) or Gbps (giga bits per second)? I see mixed usage in this comment thread so I’m left wondering what it actually is. The article is consistent and uses Gigabytes.
- gregorygoc 1y agoIt’s mind boggling why this resource has not been provided by NVIDIA yet. It reached the point that 3rd parties reverse engineer and summarize NV hardware to a point it becomes an actually useful mental model. What are the actual incentives at NVIDIA? If it’s all about marketing they’re doing great, but I have some doubts about engineering culture.
- threeducks 1y agoWith mediocre documentation, NVIDIAs closed-source libraries, such as cuBLAS and cuDNN, will remain the fastest way to perform certain tasks, thereby strengthening vendor lock-in. And of course it makes it more difficult for other companies to reverse engineer.
- hackrmn 1y agoPlenty of circumstantial evidence pointing to the fact NVIDIA prefers to hand out semi-tailored documentaion resources to signatories and other "VIPs", if not the least to exert control over who and how uses their products. I wouldn't put it past them to routinely neglect their _public_ documentation, for one reason or another that makes commercial sense to them but not the public. As for incentives, go figure indeed -- you'd think by walling off API documentation, they're shooting themselves in the feet every day, but in these days of betting it all on AI, which means selling GPUs, software and those same NDA-signed VIP-documentation articles to "partners", maybe they're all set anyway and care even less for the odd developer who wants to know how their flagship GPU works.
- KeplerBoy 1y agoNvidia has ridiculously good documentation for all of this compared to its competitors.
- dahart 1y agoWhat makes you think that? It appears most of this material came straight out of NVIDIA documentation. What do you think is missing? I just checked and found the H100 diagram for example is copied (without being correctly attributed) from the H100 whitepaper: https://resources.nvidia.com/en-us-hopper-architecture/nvidia-h100-tensor-c https://resources.nvidia.com/en-us-hopper-architecture/nvidi... Much of the info on compute and bandwidth is from that and other architecture whitepapers, as well as the CUDA C++ programming guide, which covers a lot of what this article shares, in particular chapters 5, 6, and 7. https://docs.nvidia.com/cuda/cuda-c-programming-guide/ https://docs.nvidia.com/cuda/cuda-c-programming-guide/ There’s plenty of value in third parties distilling and having short form versions, and of writing their own takes on this, but this article wouldn’t have been possible without NVIDIA’s docs, so the speculation, FUD and shade is perhaps unjustified.
- akshaydatazip 1y agoThanks for the really thorough research on that . Right what I wanted for my morning coffee
- physicsguy 1y agoIt’s interesting that nvshmem has taken off in ML because the MPI equivalents were never that satisfactory in the simulation world. Mind you, I did all long range force stuff which is difficult to work with over multiple nodes at the best of times.
- tomhow 1y agoDiscussion of original series: How to scale your model: A systems view of LLMs on TPUs - https://news.ycombinator.com/item?id=42936910 https://news.ycombinator.com/item?id=42936910 - Feb 2025 (30 comments)
- radarsat1 1y agoA comment from there: > There are plans to release a PDF version; need to fix some formatting issues + convert the animated diagrams into static images. I don't see anything on the page about it, has there been an update on this? I'd love to put this on my e-reader.
- tucnak 1y agoThis post is a great illustration why TPU's lend more nicely towards homogenous computing: yes, there's systolic array limitations (not good for sparsity) but all things considering, bandwidth doesn't change as your cluster ever so larger grows. It's a shame Google is not interested in selling this hardware: if they were available, it would open the door to compute-in-network capabilities far beyond what's currently available; by combining non-homogenous topologies involving various FPGA solutions, i.e. with Alveo V80 exposing 4x800G NIC's. Also: it's a shame Google doesn't talk about how they use TPU's outside of LLM.
- namibj 1y agoDo TPUs allow having a variable array dimension at somewhat inner nesting level of the loop structure yet? Like, where you load expensive (bandwidth-heavy) data in from HBM, process a variable-length array with this, then stow away/accumulate into a fixed-size vector? Last I looked they would require the host to synthesize a suitable instruction stream for this on-the-fly with no existing tooling to do so efficiently. An example where this would be relevant would be LLM inference prefill stage with (activated) MoA expert count on the order of — to a small integer smaller than — the prompt length, where you'd want to only load needed experts and only load each one at most once per layer.
- tormeh 1y agoI find it very hard to justify investing time into learning something that's neither open source nor has multiple interchangeable vendors. Being good at using Nvidia chips sounds a lot like being an ABAP consultant or similar to me. I realize there's a lot of money to be made in the field right now, but IIUC historically this kind of thing has not been a great move.
- saagarjha 1y agoSure, but you can make money in the field and retire faster than it becomes irrelevant. FWIW none of the ideas here are novel or nontransferable–it's just the specific design that is proprietary. Understanding how to do an AllReduce has been of theoretical interest for decades and will probably remain worth doing far into the future.
- j45 1y agoTech is always like this. You move from one thing to the next. With your transferable skills, experience and thinking that is beyond one programming language. Even Apple is simply exporting to CUDA now.
- victor106 1y ago> Even Apple is simply exporting to CUDA now. Really!!! Any resources you can share?
- j45 1y agoIt’s one way but still something. https://9to5mac.com/2025/07/15/apples-machine-learning-framework-is-getting-support-for-nvidia-gpus/ https://9to5mac.com/2025/07/15/apples-machine-learning-frame...
- almostgotcaught 1y ago> Even Apple is simply exporting to CUDA now. This is like when journalists write clickbait article titles by omitting all qualifiers (eg "states banning fluoride" when it's only some states). One framework added a CUDA backend. You think all of Apple uses only one framework? Further what makes you think this even gets internal use?
- evrennetwork 1y ago[dead]
- hackrmn 1y agoI find the piece, much like a lot of other documentation, "imprecise". Like most such efforts, it likely caters to a group of people expected to benefit from being explained what a GPU is, but it fumbles it terms, e.g. (the first image with burned-in text): > The "Warp Scheduler" is a SIMD vector unit like the TPU VPU with 32 lanes, called "CUDA Cores" It's not clear from the above what a "CUDA core" (singular) _is_ -- this is the archetypical "let me explain things to you" error most people make, in good faith usually -- if I don't know the material, and I am out to understand, then you have gotten me to read all of it but without making clear the very objects of your explanation. And so, for these kind of "compounding errors", people who the piece was likely targeted at, are none the wiser really, while those who already have a good grasp of the concepts attempted explained, like what a CUDA core actually is, already know most of what the piece is trying to explain anyway. My advice to everyone who starts out with a back of envelope cheatsheet then decides to publish it "for the good of mankind", e.g. on Github: please be surgically precise with your terms -- the terms are your trading cards, then come the verbs etc. I mean this is all writing 101, but it's a rare thing, evidently. Don't mix and match terms, don't conflate them (the reader will do it for you many times over for free if you're sloppy), and be diligent with analogies. Evidently, the piece may have been written to help those already familiar with TPU terminology -- it mentions "MXU" but there's no telling what that is. I understand I am asking for a tall order, but the piece is long and all the effort that was put in, could have been complemented with minimal extra hypertext, like annotated abbreviations like "MXU". I can always ask $AI to do the equivalent for me, which is a tragedy according to some.
- einpoklum 1y ago> It's not clear from the above what a "CUDA core" (singular) _is_ A CUDA core is basically a SIMD lane on an actual core on an NVIDIA GPUs. For a longer version of this answer: https://stackoverflow.com/a/48130362/1593077 https://stackoverflow.com/a/48130362/1593077
- pklausler 1y agoSo it's a "SIMD lane" that can itself perform actual SIMD instructions? I think you want a metaphor that doesn't also depend on its literal meaning.
- einpoklum 1y agoWe should remember that these structural diagrams are _not_ necessarily what NVIDIA actually has as hardware. They carefully avoid guaranteeing that any of the entities or blocks you see in the diagrams actually _exist_. It is still just a mental model NVIDIA offers for us to think about their GPUs, and more specifically the SMs, rather than a simplified circuit layout. For example, we don't know how many actual functional units an SM has; we don't know if the "tensor core" even _exists_ as a piece of hardware, or whether there's just some kind of orchestration of other functional units; and IIRC we don't know what exactly happens at the sub-warp level w.r.t. issuing and such.
- KeplerBoy 1y agoInteresting perspective. Aren't SMs basically blocked while running tensor core operations, which might hint that it's the same FPUs doing the work after all?
- einpoklum 1y agoI doubt that can fully be the case, because there are other functional units on SMs, like Load/Store, ALU / Integer ops, and Special Function Units. But you may be right, we would need to consult the academic "investigatory" papers or blog posts and see whether this has been checked.
- radarsat1 1y agoWhy haven't Nvidia developed a TPU yet?
- Philpax 1y agoThey don't need to. Their hardware and programming model are already dominant, and TPUs are harder to program for.
- dist-epoch 1y agoThis article suggests they sort of did: 90% of the flops is in matrix multiplication units. They leave some performance on the table, but they gain flexible compilers.
- HarHarVeryFunny 1y agoMeaning what? Something less flexible? Less CUDA cores and more Tensor Cores? The majority of NVidia's profits (almost 90%) do come from data center, most of which is going to be neural net acceleration, and I'd have to assume that they have optimized their data center products to maximize performance for typical customer workloads. I'm sure that Microsoft would provide feedback to Nvidia if they felt changes were needed to better compete with Google in the cloud compute market.
- cwmoore 1y ago> most of which is going to be neural net acceleration is it?
- HarHarVeryFunny 1y agoI've got to assume so, since data center revenue growth seems to have grown in sync with recent growth in AI adoption. CUDA has been around for a long time, so it would seem highly coincidental if non-AI CUDA usage was only just now surging at same time as AI usage is taking off, and new data center build announcements seem to invariably be linked to AI.
- business_liveit 1y agoso, Why didn't Nvidia developed a TPU yet?
- cwmoore 1y agoProbably proprietary. Go GOOG. I like how bad your comment is.
- varelse 1y ago[dead]
- pbrumm 1y agoIf you have optimized your math heavy code and it is already in a typed language and you need it to be faster, then you think about the GPU options In my experience you can roughly get 8x speed improvement. Turning a 4 second web response into half a second can be game changing. But it is a lot easier to use a web socket and put a spinner or cache result in background. Running a GPU in the cloud is expensive
- gchadwick 1y agoThis whole series is fantastic! Does an excellent job of explaining the theoretical limits to running modern AI workloads and explains the architecture and techniques (in particular methods of parallelism) you can use. Yes it's all TPU focussed (other than this most recent part) but a lot of what it discusses are generally principles you can apply elsewhere (or easy enough to see how you could generalise them).
- ngcc_hk 1y agoThis is part 12 … the title seems to hint on how do one think about Gpu today … eg why llm comes about. Instead it is about cf with tpu? And then I note the part 12 … not sure what one should expect to jump in the middle of a whole series and what … well may stop and move on.
- boxerab 1y ago"How to Think About NVIDIA GPUS" is a better title
- aktuel 1y agoWhat is the "Use completions" toggle supposed to do? If I enable it I just get empty responses.