4 ms·Gemlite: Towards Building Custom Low-Bit Fused CUDA Kernels47 points by un_ess 2y agojohnsutor 2y agoThis would be great to have for the Triton language as well.apsec112 2y agoWeird that they don't mention Triton? I only skimmed it, but I'm not sure what the pros and cons would be vs. Triton, which is the tool I'd use if I wanted custom quantized inference kernels.