4 ms·
Yes, when you use the PTX backend it supports Tensor Cores.It has also implementation for flash attention. You can also write your own kernels, have a look here
by mikepapadim 10mo ago
Yes, when you use the PTX backend it supports Tensor Cores.It has also implementation for flash attention. You can also write your own kernels, have a look here: https://github.com/beehive-lab/GPULlama3.java/blob/main/src/main/java/org/beehive/gpullama3/tornadovm/kernels/TransformerComputeKernels.java https://github.com/beehive-lab/GPULlama3.java/blob/main/src/... https://github.com/beehive-lab/GPULlama3.java/blob/main/src/main/java/org/beehive/gpullama3/tornadovm/kernels/TransformerComputeKernelsLayered.java https://github.com/beehive-lab/GPULlama3.java/blob/main/src/...
- lostmsu 10mo agoTornadoVM GitHub has no mentions of tensor cores or WMMA instructions. The only mention of tensor cores is in 2024 and states they are not used: https://github.com/beehive-lab/TornadoVM/discussions/393 https://github.com/beehive-lab/TornadoVM/discussions/393
- mikepapadim 10mo agohttps://github.com/beehive-lab/TornadoVM/pull/732 https://github.com/beehive-lab/TornadoVM/pull/732 https://github.com/beehive-lab/TornadoVM/pull/313 https://github.com/beehive-lab/TornadoVM/pull/313
- lostmsu 10mo agoI believe these are SIMD. Tensor cores require MMA family of instructions. Ask me how I know. :) https://github.com/m4rs-mt/ILGPU/compare/master...lostmsu:ILGPU:wgmma/make-descriptor https://github.com/m4rs-mt/ILGPU/compare/master...lostmsu:IL... Good article: https://alexarmbr.github.io/2024/08/10/How-To-Write-A-Fast-Matrix-Multiplication-From-Scratch-With-Tensor-Cores.html#tensor-core-vs-ffma https://alexarmbr.github.io/2024/08/10/How-To-Write-A-Fast-M...