4 ms·
I see the term 'local inference' everywhere. It's an absurd misnomer without hardware and cost defined. I can also run a coal fired power plant in my backyard,
by aliljet 1y ago
I see the term 'local inference' everywhere. It's an absurd misnomer without hardware and cost defined. I can also run a coal fired power plant in my backyard, but in practice, there's no reasonable way to make that economical beyond being a toy.
(And I should add, you are a hero for doing this work, only love in my comment, but still a demand for detail$!)
- regularfry 1y agoHardware and cost is assumed to be approximately desktop-class. If you've got a gaming rig with an RTX 4090 and 128MB RAM, you can run this if you pick the right quant.
- cmpxchg8b 1y ago128MB? Quantization has come a long way!
- danielhanchen 1y agoI think they mis-spoke 128GB* :)
- regularfry 1y agoWishful thinking there on my part.
- danielhanchen 1y agoThough technically < 1GB is enough - one had offload it to the SSD, albeit with very slow speeds!
- danielhanchen 1y agoThe trick of llama.cpp and our dynamic quants is you can actually offload the model to RAM / even an SSD! If you have GPU VRAM + RAM + SSD > the model size (say 90GB for dynamic 2bit quant), then it'll run well! Ie you can actually run it on a local desktop or even your laptop now! You don't need a 90GB GPU for example, but say a 24GB GPU + 64GB to 128GB RAM. The speeds are around 3 to 5 tokens / second, so still ok! I write more about improving speed for local devices here: https://docs.unsloth.ai/basics/qwen3-how-to-run-and-fine-tune/qwen3-2507#improving-generation-speed https://docs.unsloth.ai/basics/qwen3-how-to-run-and-fine-tun...