4 ms·
This holds half-true. What we decided to do is essentially just hand-rolling PTX with the agents under a harness. If there is lacking some function, we write/t
by augment_me 2mo ago
This holds half-true.
What we decided to do is essentially just hand-rolling PTX with the agents under a harness. If there is lacking some function, we write/take a CuteDSL kernel, export PTX on compilation, transform it with the harness and then feed it into the agents.
We beat cuBLAS, and we beat all of abstractions - Triton, Helion, TLX, Mojo, ThunderKittens, cuteDSL, etc on Hopper and Blackwell for all of our workloads over 3-4 weeks.
vLLM dropped torch.compile support because they realized that they programmers were good enough to just generate the Triton kernels directly for all the passes efficiently.
If you work with this for prod the writing is on the wall sadly. The abstraction layer is just really much lower if you want full perf.
- vatsachak 2mo agoBut you end up writing your own abstractions in the process?
- augment_me 2mo agoIsh, today I am on the fence because sometimes(half-half) its necessary, a year ago it was completely impossible, but I think its a question of context length, attention and task. Modern models can handle raw PTX pretty well. If you have below 20k lines, it's able to make its own abstractions in the thinking loop Nth pass when it needs to. The only reason we would have to make our own abstractions today would be to fill the limits of the current context context window attention trade-offs the companies make. I think the article is calling for the redundancy of high-level tiling abstractions. There is simply no need to work at a higher level and give up performance today when code is free.
- vatsachak 2mo agoI feel like both abstractions and LLMs are time savers, so why not combine both. Also, if I don't create abstractions and guide the LLM I've found that the codebase starts getting out of my reach. And until LLMs can maintain code ad infinitum I like to understand my code. I can see why it might not matter in hardcore performance coding though.
- augment_me 2mo agoIf the task is to build something that requires abstractions, sure. But the article in question is in GPU-optimization-land, which happens to be an extremely verifiable and hill-climbable domain. Look at recent GPU programming competitions and find that top-10 solutions are completely generated. In general the field of compilers is in a crisis because why wait for compilations when you can just generate the end result and hill- climb it. I am kind of in the middle of this and I see it going toward LLM-hill-climbers by the day.