4 ms·
There are companies whose whole job right now is to optimize kernels so that things run faster. I wonder if those companies are going to be dethroned by some so
by orliesaurus 3mo ago
There are companies whose whole job right now is to optimize kernels so that things run faster. I wonder if those companies are going to be dethroned by some sort of like open source library that can do that really well (I bet Nvidia could release it any day.).. or if they're going to thrive and be acquired by the big providers as a `moat` to speed up their infrerence.
- spmurrayzzz 3mo agoNear-term acquihires are certainly a likely bet I think. But given model progress on related benchmarks like kernelbench [1], I do think a set of more commoditized solutions is also inevitable. The caveat though is that each new gen of hardware often comes with brand new constraints/features that a given generation of models haven't seen before (e.g. tcgen05 in blackwell was OOD at one point). As the models start to generalize better, this might not be a showstopper, but still an issue at least currently. [1] https://kernelbench.com/ https://kernelbench.com/
- xpct 3mo agoThank you for sharing the link. It's fascinating that the models can only do 10-20% on the hard subset, and I wonder why that is so. The fact that they can only get 30-40% out of the fp8 GEMM seems unintuitive to me, I would've expected a convergence near ~80%.
- spmurrayzzz 3mo agoI'm not entirely up to date with the latest batch, but I've reviewed some of the rollouts in the past and my sense is that the models are surprisingly good at getting correct custom kernels in the happy path, but still weak at sustained/shape-robust workloads. Having to deal with writing the full path from scratch compounded by weird memory layouts, odd sizes, routing, unpacking quantized weights, etc. is definitely challenging. Also, at least a portion of this you could argue is arbitrary and entirely scoped to the eval itself. The fp8 GEMM score could be low simply because one of the shapes is fairly skinny (i.e. not enough math work to keep the compute engine busy for a meaningful amount of time).
- connicpu 3mo agoWhen you run CUDA at scale dealing with nvidia driver and library bugs takes up a disgustingly large percentage of engineer time, I don't know a lot of people who would be looking forward to rely on more nvidia libraries.
- orliesaurus 3mo agofair point, but are there alternatives that aren't CUDA locked?
- whattheheckheck 3mo agoIs there an issue board for these bugs? I want to see what is a disgustingly large percent. 50%?
- david-gpu 3mo agoHow do you determine that the bugs you run into are located in the Nvidia drivers and libraries? Way back when I wrote the OpenCL driver at Qualcomm, we would frequently get bug reports from customers complaining about our code. During my tenure, every single one of them was root-caused as an application bug. Unsurprisingly, considering that our code was backed by an extensive test suite and their code wasn't. Not to say that our code was perfect, of course. But people have a tendency to blame GPU drivers when the problem often lies elsewhere.
- connicpu 3mo agoIf you're big enough you can get direct access to Nvidia engineers, and they are usually transparent when they find out the bug was in their software and send you a patched version to try to resolve the issue
- AlotOfReading 3mo agoAnd when the bug is in hardware instead, good luck getting them to admit it. Instead they go "We can't confirm that there's an issue, but it'll be gone in the next revision and meanwhile don't use that instruction sequence".
- einpoklum 3mo agoProbably not, because the specifics of the workload - exact parameters, representation of data in memory, value ranges etc - lead you to highly divergent optimization strategies.
- orliesaurus 3mo agoshouldn't it be possible to be run as a mlautoresearch project? i.e. orchestrate 10 strategies to speed it up, run in paralellel, pick the winning and go from there?
- einpoklum 3mo agoYou are assuming all problems in the world are solvable by one of "10 strategies".
- charcircuit 3mo agoHe is not assuming that as he includes "and go from there."
- saagarjha 3mo agoNo, because all the low hanging fruit that this kind of thing would find has usually been picked.
- keynha 3mo ago[flagged]