3 ms·
Multiply FP8 matrices with FP32 scaling factors giving a bfloat16 matrix result on an nVidia Hopper or newer GPU.
by devit 2y ago
Multiply FP8 matrices with FP32 scaling factors giving a bfloat16 matrix result on an nVidia Hopper or newer GPU.
- Maxious 2y agoJust tested and it doesn't work out of the box on the consumer 50 series ie. 5080: Testing GEMM: Assertion failed: deep_gemm/jit/../include/deep_gemm/fp8_gemm.cuh:369, condition: cudaFuncSetAttribute(kernel, cudaFuncAttributeMaxDynamicSharedMemorySize, smem_size) == cudaSuccess terminate called after throwing an instance of 'AssertionException' what(): Assertion failed: cudaFuncSetAttribute(kernel, cudaFuncAttributeMaxDynamicSharedMemorySize, smem_size) == cudaSuccess
- deleted 2y ago[deleted]
- wenc 2y agoIt says: > DeepGEMM exclusively supports NVIDIA Hopper tensor cores
- devit 2y agoPerhaps your card has less per-SM shared memory than the GPUs DeepSeek uses. Try to lower the sm90_capacity value in gemm.py: I think 128KB is the correct value for RTX 5080 compared to 256KB for the H100/H800. And probably add ", 3, 2, 1" after "6, 5, 4".