3 ms·
I am personally a really huge fan of cutlass, and I've almost read every single file in their `include/cutlass/` folder (haven't followed up with the `cute` stu
by junrushao1994 4y ago
I am personally a really huge fan of cutlass, and I've almost read every single file in their `include/cutlass/` folder (haven't followed up with the `cute` stuff yet).
Just like you said, really appreciate that we could actually understand what is going on internally inside the kernel with cutlass, and customize it in a way that cuBLAS doesn't necessarily provide.
Have to agree with you that the template stuff is really annoying. Well, even with some template tricks, the error messages are still less readable, and it is where I think a better abstraction could benefit. Imagine you have an abstraction that simplifies the those threadblocks/warps/etc in a unified way while generalizing it to more backends (AMDGPU, Vulkan, AVX512-VNNI, etc), providing more friendly error messages along compilation given the abstraction is almost certainly more structured than pure c++ code.
- touisteur 4y agoI know what you mean. I guess Triton might be one way people are actually trying that, and further than that there's a lot of people putting years of work into MLIR-based tech, trying to abstract in a retargetable way the algorithm from its 'scheduling'. Might be worth a look if you're into this :-)
- junrushao1994 4y agoAh nice! I intentionally didn’t talk a lot about “scheduling” because 1) I’m personally heavily working on it, which potentially makes a conflict of interest, and 2) I don’t want to deviate a lot from the topic in this thread about “optimize matmuls”. Check out my google scholar for more details though! I love MLIR and Modular, so please do share more about it! If it’s potential distraction from this thread, I’m also open to email communication if you are interested! Oh btw, to clarify, I’m not saying Triton is an ideal abstraction. I love it and it’s super popular because it’s the most user-friendly option for ML researchers to write performant kernels on certain gpus, but from a MLSys researcher’s perspective, I’m personally more ambitious and wanted to target broader range of hardwares. Also I really appreciate Philippe’s work that makes Triton really performant and easy to use.
- deleted 4y ago[deleted]