3 ms·
Nice! His Shakespeare generator was one of the first projects I tried after ollama. The goal was to understand what LLMs were about. I have been on an LLM bing
by sieve 1y ago
Nice! His Shakespeare generator was one of the first projects I tried after ollama. The goal was to understand what LLMs were about.
I have been on an LLM binge this last week or so trying to build a from-scratch training and inference system with two back ends:
- CPU (backed by JAX)
- GPU (backed by wgpu-py). This is critical for me as I am unwilling to deal with the nonsense that is rocm/pytorch. Vulkan works for me. That is what I use with llama-cpp.
I got both back ends working last week, but the GPU back end was buggy. So the week has been about fixing bugs, refactoring the WGSL code, making things more efficient.
I am using LLMs extensively in this process and they have been a revelation. Use a nice refactoring prompt and they are able to fix things one by one resulting in something fully functional and type-checked by astral ty.
- danielmarkbruce 1y agoUnwilling to deal with pytorch? You couldn't possibly hobble yourself anymore if you tried.
- sieve 1y agoIf you want to train/sample large models, then use what the rest of the industry uses. My use case is different. I want something that I can run quickly on one GPU without worrying about whether it is supported or not. I am interested in convenience, not in squeezing out the last bit of performance from a card.
- danielmarkbruce 1y agoYou wildly misunderstand pytorch.
- sieve 1y agoWhat is there to misunderstand? It doesn't even install properly most of the time on my machine. You have to use a specific python version. I gave up on all tools that depend on it for inference. llama-cpp compiles cleanly on my system for Vulkan. I want the same simplicity to test model training.
- danielmarkbruce 1y agopytorch is as easy as you are going to find for your exact use case. If you can't handle the requirement of a specific version of python, you are going to struggle in software land. ChatGPT can show you the way.
- sieve 1y agoI have been doing this for 25 years and no longer have the patience to deal with stuff like this. I am never going to install Arch from scratch by building the configuration by hand ever again. The same with pytorch and rocm. Getting them to work and recognize my GPU without passing arcane flags was a problem. I could at least avoid the pain with llama-cpp because of its vulkan support. pytorch apparently doesn't have a vulkan backend. So I decided to roll out my own wgpu-py one.
- danielmarkbruce 1y agoFair enough I guess. I think you'll find the relatively minor headache worth it. Pytorch brings a lot to the table.
- rpdillon 1y agoFWIW, I've been experimenting with LLMs for the last couple of years, and have exclusively built everything I do around llama.cpp exactly because of the issues you highlight. "gem install hairball" has gone way too far, and I appreciate shallow dependency stacks.
- nl 1y agoI suspect the OP's issues might be mostly related to the ROCM version of PyTorch. AMD still can't get this right.
- danielmarkbruce 1y agoProbably - but the answer is to avoid ROCM, not pytorch.
- yorwba 1y agoAvoiding ROCm means buying a new Nvidia GPU. Some people would like to keep using the hardware they already have.
- danielmarkbruce 1y agoThe cost to deal with rocm is > cost of a consumer nvidia gpu by orders of magnitude.
- ComputerGuru 1y agoIf you’re not writing/modifying the model itself but only training, fine tuning, and inferencing, ONNX now supports these with basically any backend execution provider without needing to get into dependency version hell.
- Breza 1y agoWhat are your thoughts on using JAX? I've used TensorFlow and Pytorch and I feel like I'm missing out by not having experience with JAX. But at the same time, I'm not sure what the advantages are.
- sieve 1y agoI only used it to build the CPU back end. It was a fair bit faster than the previous numpy back end. One good thing about JAX (unlike numpy) is that it also gives you access to a GPU back end if you have the appropriate stuff installed.