4 ms·
Super curious how you did this! Doesn’t 30B model require a hefty computer to run locally (assuming you’re tuning a non-quantized version)
by syntaxing 3y ago
Super curious how you did this! Doesn’t 30B model require a hefty computer to run locally (assuming you’re tuning a non-quantized version)
- drdaeman 3y agoNot at all. Even a Raspberry Pi would do - you only need ~6GiB RAM for a 4-bit quantized LLaMA model (though it's gonna be quite slow). A decent modern desktop machine would do just fine, no need for anything extra fancy. What I'm wondering is how they fed the documents, as all those LLMs have limitations on the input sizes.
- yyyk 3y ago>~6GiB RAM That's for the 7B model. The 30B model needs 24GB quantized (or 64GB for the unquantized model).
- jstarfish 3y agoI've seen reports that it wrecks RPi SD cards in short order though, so beware... > What I'm wondering is how they fed the documents, as all those LLMs have limitations on the input sizes. It's like file hashing at scale, you don't have to read the whole stream for every file, just the first 1024/2048 bytes (or first few paragraphs). (This works for classification and sorting, less so for summarization.)
- yacine_ 3y agoI'm running 30b quant on my 3090. that many 4bits fortunately fit into my precious vram