10 ms·
We bought the whole GPU, so we're damn well going to use the whole GPU
- matthewmorgan 1y ago[flagged]
- Waterluvian 1y agoKind of. We like to call them grad students.
- wizzwizz4 1y agoThey're cheaper!
- lemonlearnings 1y agoAs long as they damn well use their whole brain.
- discordance 1y agoTime to first token is 20 odd years and 100k of education debt
- gpm 1y agoYeah, but the previous "investors" ate up that cost and left pure profit for us "founders" (professors).
- threemux 1y agoWe've finally achieved AGI* ! * Actually Graduate Individuals
- cscheid 1y agoScience is just Stochastic Graduate Descent, as we used to say.
- QQ00 1y agoAI = Academic Interns
- jonstewart 1y agoFigure 1: Zoooommmm Accept!
- barrkel 1y agoI'm reminded of how Carmack talked about the extra efficiencies available when targeting consoles, because you knew exactly what hardware was available. It's great that the efficiencies available can be shown to be extractable. The real, much harder, trick is putting together a sufficiently smart compiler to enable them for heterogeneous compute setups.
- deleted 1y ago[deleted]
- bombcar 1y agoThe demoscene also is an example of how much you can do if you can be absolutely sure exactly what hardware you’re running on. The problem is that even for things like consoles, it's usually more "cost efficient" to write normal fast-to-write code that isn't maximally effective, let the compiler do its magic, and call it good enough. Sometimes I dream of what the world would do if we were mystically stuck on exactly the processors we have today, for twenty years.
- eru 1y agoIt's not just being sure exactly what the hardware is, in demos you have the additional luxury of not being interactive. So you can plan everything exactly out in advance.
- saagarjha 1y agoThis is true of inference too.
- potatolicious 1y agoConsoles are pretty heterogeneous IRL too, though. You have multiple SKUs (regular and Pro, for example), not to mention most games will also target multiple consoles (PlayStation + Xbox + Switch is a common combo). So in reality the opportunities to really code against a specific piece of hardware are few and far between... Heck, then you get into multiple operating modes of the same hardware - the Nintendo Switch has a different perf profile if it's docked vs. not.
- Stratoscope 1y agoI know I'm being unfair, but something about the writing style reminds me of this classic: Transgressing the Boundaries: Towards a Transformative Hermeneutics of Quantum Gravity https://physics.nyu.edu/faculty/sokal/transgress_v2/transgress_v2_singlefile.html https://physics.nyu.edu/faculty/sokal/transgress_v2/transgre...
- CamperBob2 1y agoOnly a matter of time until we start seeing bogus Hard Science papers like that, now that we've given the Social Text people the tools they need to take their revenge. They will argue that we had it coming, and that it serves us right, and maybe they're not wrong.
- rcxdude 1y agoAlready been done: https://www.nationalgeographic.com/pages/article/131003-bohannon-science-spoof-open-access-peer-review-cancer https://www.nationalgeographic.com/pages/article/131003-boha...
- sydriax 1y agoBen here -- you may be amused to know that Alan Sokal was my dad's freshman roommate in undergrad!
- Stratoscope 1y agoAwesome! We truly have Transgressed the Boundaries. (And I'm curious... The way you said "Ben here" makes me wonder if I know you?)
- sciurus 1y ago(I think we was just introducing himself as Ben Spector, the lead author of the paper.)
- 1y ago
- aeon_ai 1y ago> It is sensitive to compiler versions, GPU setup, and sometimes even being looked at the wrong way, and we have no intention whatsoever of supporting it. My favorite type of code
- versteegen 1y agoExcellent writeup. I like the interpreter. But I can only assume all these ideas have been widely implemented at all significant labs for years, so I'm surprised to see this written in 2025. This is all about taking things to their logical conclusions, not arcane magic. If you're going to spend billions on GPUs, why wouldn't you spend a little on CUDA programmer hours?
- sailingparrot 1y ago> I can only assume all these ideas have been widely implemented at all significant labs for years, right? Nope. I was also surprised, when joining such significant labs at how much relatively-low hanging fruits were still available to work on. But the reality is that there is just too much work to do, each seemingly super-important, and not enough people to do it.
- versteegen 1y agoI'm very willing to believe that. When I hear that they just don't have enough staff for it I get the impression is that they set their hiring bar for engineers too high. Optimising CUDA is quite different from having experience training LLMs.
- sailingparrot 1y ago> they set their hiring bar for engineers too high Not sure I agree, if you look at the head count growth of companies like OpenAI, Anthropic etc, it is super fast, its already pretty hard to keep everything working smoothly with that rate of employee growth, so going faster than that seems very risky. Ultimately I think it's mostly caused by the field still being so new. Everything still needs to be optimized and there just aren't that many very good CUDA programmers to start with, then you need to find one that also has deep knowledge of ML and transformers architectures, which further drains the pool. And then when you do find one of them, there is 50 different things they could be working on instead of what's in the article, all equally or more impactful. The architectures being constantly evolving also make it hard/not a great ROI to go super super deep in single digit % optimization when there is new stuff coming out all the time that can be made an order of magnitude faster. A good example of that is flash attention: it is maybe the most significant/impactful optimization in ML of the last few years. Tl;dr is how do you fuse the entire attention pipeline together to make it much faster and avoid massive tensor materialization. The bottleneck was obvious to anyone that profiled a Transformer-based model, but there was no obvious solution because of how softmax works. Yet the paper that ultimately unblock this was published back in 2019 [1], but it took 3 years for a team to connect the dots. Most people in pure ML engineering didn't know about the paper and don't have good enough CUDA knowledge/ GPU arch understanding, most people with good CUDA knowledge don't understand ML well enough, and even the author of that 2019 paper said "[we] hypothesize that this reduction in memory accesses should improve Softmax performance on actual hardware" but didn't have the technical skills to test this or to see how that could be part of a bigger breakthrough because it requires understanding core concepts in how GPU worked and compute/memory imbalance. [1]: https://arxiv.org/pdf/1805.02867 https://arxiv.org/pdf/1805.02867
- throw0101d 1y agoIf your workload can't actually use the whole (NVidia) GPU, it is possible to slice it up so that it can be shared between multiple users: * https://docs.nvidia.com/datacenter/tesla/mig-user-guide/ https://docs.nvidia.com/datacenter/tesla/mig-user-guide/ * https://www.nvidia.com/en-us/technologies/multi-instance-gpu/ https://www.nvidia.com/en-us/technologies/multi-instance-gpu... Or having multiple processes from one user share it: * https://docs.nvidia.com/deploy/mps/index.html https://docs.nvidia.com/deploy/mps/index.html
- jsheard 1y agoAIUI only on workstation/server cards though, it's one of the levers they pull to artificially segment their lineup.
- bix6 1y agoHow real is the risk of information leakage if I’m on a shared GPU with multiple users?
- LPisGood 1y agoI remember a few years ago my hardware security professor suggested we try to implement Rowhammer on GPU. I ended up doing something else, but it looks like someone got there: https://arxiv.org/abs/2507.08166 https://arxiv.org/abs/2507.08166
- woadwarrior01 1y agoVery real. https://www.usenix.org/system/files/usenixsecurity24-guo-yanan.pdf https://www.usenix.org/system/files/usenixsecurity24-guo-yan... https://www.sciencedirect.com/science/article/pii/S0167404820303886 https://www.sciencedirect.com/science/article/pii/S016740482...
- throw0101d 1y agoI do not see MIG mentioned in either paper. I do not think the papers are examining isolation security between instances, which the GP was asking about.
- Archit3ch 1y agoThe sentiment in the title resonates, but for consumer GPUs (the article is about server cards). The recently leaked M5 benchmarks reveal a 35% faster GPU. These improvements compound, so you can get a GPU that's effectively twice as fast by waiting a couple of years. Modern GPUs are the equivalent of local supercomputers, but the drivers, languages and libraries are still playing catch up. Imagine the audio processing you could do if only you could target that hardware.
- QuantumNoodle 1y ago> Imagine the audio processing you could do if only you could target that hardware. That's an interesting thought. Commercial grade signal processing rely on FPGAs and the Fintech field adapted them for high frequency trading. I wonder if we will see signal processing enabled on GPUs for consumers if the GPU drivers were more open.
- Agentlien 1y agoIt should definitely be possible already using CUDA or computer shaders. From a theoretical view computer graphics is signal processing but with a signal consisting of up to four color channels across two dimensions. This is the view taken in a lot of papers and practical implementations. After all, a lot of computer graphics is about applying filters (post-processing) such as color grading, anti-aliasing, etc. to this signal. So, in a very real sense, signal processing is exactly what the GPU is built for and primarily used for.
- djmips 1y agoI recall there have been some efforts in audio DSP on GPUs, the audio bandwidth is low so even transporting the results back to the CPU to be played could be done fast enough to maintain a usable latency.
- bigyabai 1y agoApple gives developers almost all the compute drivers you could want from them. If you can't express your GPU acceleration as a Metal Compute Shader, you probably aren't leaving any GPU horsepower on the table. ANE and MLX will get exposed in higher-level CoreML frameworks, everyone should be happy. 35% raster improvements, it's worth noting, is not super impressive on the GPU side of things. Most raster compute is a square function, to double your render resolution you need a 4x the GPU power (on-paper) to handle the pixel count. That's what, six years of annual iteration? A large component of Apple and AMD's inability to break into Nvidia's CUDA empire is their obsession over raster optimization in a world where DLSS and FSR exists. It's a noble pursuit, but even as a gamer I've gotta admit they're wasting their time. We have software methods that can close the gap in render quality between $100 GPUs and $1000 GPUs, but no such solution for GPGPU compute.
- tuna74 1y agoIt would be nice if the article header would actually be clear that they are optimizing a CUDA chip. There is a difference between a GPU and a CUDA chip.
- KeplerBoy 1y agoYou could do similar stuff on AMD chips.
- AaronFriel 1y agoNot using the NVDEC and NVJPG units to decompress weights into registers? And you say you're using the whole GPU. There are entire blocks on the silicon going idle!
- refibrillator 1y agoHa made me chuckle. For those wondering seriously about this, it’s not a viable optimization because weights are not readily compressible via JPEG/DCT, and there are a limited number of these units on the chip which bottlenecks throughout, meaning speed is dwarfed by simply reading uncompressed weights from HBM.
- moralestapia 1y agoYeah, but they could be. I won an GPU hackathon back in 2019 doing something very similar to this; although the other way around, I was compressing weights using hardware modules.
- heavyset_go 1y agoHave a link to this?
- moralestapia 1y agoUnfortunately no. I have cool picture, though!
- heavyset_go 1y agoI will have to settle for a picture then :)
- moralestapia 1y agoSend email (see profile), I'll gladly share more details ^^.
- hinkley 1y agoI bought a car with side impact airbags, so we’re damn well going to use the side impact airbags. Maybe… you don’t actually want or need to use all the features of something you bought. Particularly given that GPUs previously used for cryptocurrency mining may have damaged themselves while being run full out for a year straight.
- winwang 1y agoI thought I was going to see something crazy like using RT cores in parallel with tensor cores. Like compiling matmul into triangle intersections.
- saagarjha 1y agoI don’t think datacenter GPUs have many of those.
- rldjbpin 1y ago> please be warned that this really is research code; it is sensitive to compiler versions, GPU setup, and sometimes even being looked at the wrong way the writeup is a classic example of what we lose through abstraction and how writing custom (and optimized) code still beats sticking to high-level implementations. i would go further and say that the "megakernel" written as part of the optimization is highly-model dependent as well. the whole "cuda moat" is from the generic implementations of the moving parts of the model architecture. at the same time, you lose a lot of performance through the generic code. it is like comparing writing a stock trading algo in next.js vs assembly. training models is another landscape altogether, so props to those who can quickly adapt to the hardware they got.