3 ms·
Is this chasing impossible - not criticize, and love effort- ? But is it -a little- really possible to run an LLM in a single machine ? I want to believe :)
by firatsarlar 3y ago
Is this chasing impossible - not criticize, and love effort- ?
But is it -a little- really possible to run an LLM in a single machine ?
I want to believe :)
- ranguna 3y agoI don't get where your question is coming from, you can already run LLMs on a single machine. Checkout llama.cpp, tabby, text generation webui, gpt4all, AI Dungeon open source models like clover-edition, and know this we gpu based app.
- firatsarlar 3y agoThe question comes from a kind of confusion. We know the requirements of LLMs. How can we run the hardware it is currently working on, only the big LLM, with an 11Gb graphics card? I really didn't mind!
- quickthrower2 3y agoI just tried it and it works. And works amazing compared to anything that existed anywhere on earth one school term ago! so yeah why not?
- Tepix 3y agoGTPQ has been the missing piece, it allows quantizing the model weights from 16 to 4 bits with only a small loss in quality. That it turn allows running even the large 65 billion parameter version of the LLaMA model in ~33GB of RAM or VRAM. With VRAM that requires two 24GB GPUs which is no longer completely out of reach. The model running in the browser is a smaller version with 7 billion parameters, which is good enough for some things.