4 ms·
Show HN: GPULlama3.java Llama Compilied to PTX/OpenCL Now Integrated in Quarkus
wget https://github.com/beehive-lab/TornadoVM/releases/download/v2.1.0/tornadovm-2.1.0-opencl-linux-amd64.zip https://github.com/beehive-lab/TornadoVM/releases/download/v...
unzip tornadovm-2.1.0-opencl-linux-amd64.zip
# Replace <path-to-sdk> manually with the absolute path of the extracted folder
export TORNADO_SDK="<path-to-sdk>/tornadovm-2.1.0-opencl"
export PATH=$TORNADO_SDK/bin:$PATH
tornado --devices
tornado --version
# Navigate to the project directory
cd GPULlama3.java
# Source the project-specific environment paths -> this will ensure the
source set_paths
# Build the project using Maven (skip tests for faster build)
# mvn clean package -DskipTests or just make
make
# Run the model (make sure you have downloaded the model file first - see below)
./llama-tornado --gpu --verbose-init --opencl --model beehive-llama-3.2-1b-instruct-fp16.gguf --prompt "tell me a joke"
- mikepapadim 10mo agohttps://github.com/beehive-lab/GPULlama3.java https://github.com/beehive-lab/GPULlama3.java
- lostmsu 10mo agoDoes it support flash attention? Use tensor cores? Can I write custom kernels? UPD. found no evidence that it supports tensor cores, so it's going to be many times slower than implementations that do.
- mikepapadim 10mo agoYes, when you use the PTX backend it supports Tensor Cores.It has also implementation for flash attention. You can also write your own kernels, have a look here: https://github.com/beehive-lab/GPULlama3.java/blob/main/src/main/java/org/beehive/gpullama3/tornadovm/kernels/TransformerComputeKernels.java https://github.com/beehive-lab/GPULlama3.java/blob/main/src/... https://github.com/beehive-lab/GPULlama3.java/blob/main/src/main/java/org/beehive/gpullama3/tornadovm/kernels/TransformerComputeKernelsLayered.java https://github.com/beehive-lab/GPULlama3.java/blob/main/src/...
- lostmsu 10mo agoTornadoVM GitHub has no mentions of tensor cores or WMMA instructions. The only mention of tensor cores is in 2024 and states they are not used: https://github.com/beehive-lab/TornadoVM/discussions/393 https://github.com/beehive-lab/TornadoVM/discussions/393
- mikepapadim 10mo agohttps://github.com/beehive-lab/TornadoVM/pull/732 https://github.com/beehive-lab/TornadoVM/pull/732 https://github.com/beehive-lab/TornadoVM/pull/313 https://github.com/beehive-lab/TornadoVM/pull/313
- lostmsu 10mo agoI believe these are SIMD. Tensor cores require MMA family of instructions. Ask me how I know. :) https://github.com/m4rs-mt/ILGPU/compare/master...lostmsu:ILGPU:wgmma/make-descriptor https://github.com/m4rs-mt/ILGPU/compare/master...lostmsu:IL... Good article: https://alexarmbr.github.io/2024/08/10/How-To-Write-A-Fast-Matrix-Multiplication-From-Scratch-With-Tensor-Cores.html#tensor-core-vs-ffma https://alexarmbr.github.io/2024/08/10/How-To-Write-A-Fast-M...
- sliicemasternet 10mo ago[dead]