4 ms·
Talk on optimizing matrix multiplication with Triton kernels, focusing on low-bit processing and efficient quantization for high-performance AI models.
by ibuildthings 2y ago
Talk on optimizing matrix multiplication with Triton kernels, focusing on low-bit processing and efficient quantization for high-performance AI models.