4 ms·
That’s 68 billions of parameters. It probably does not fit on ram. Though If you encode each parameter using one byte, you would need 68GB RAM which you could g
by junipertea 4y ago
That’s 68 billions of parameters. It probably does not fit on ram. Though If you encode each parameter using one byte, you would need 68GB RAM which you could get on workstations at this point.
- taf2 4y agoSeems to use about 40~ GB RAM here...
- terafo 4y agoIt fits, whisper.cpp uses 4 bit quantization, 13B model takes a little bit more than 8gb and around 9gb ram while inferencing.
- gymbeaux 4y agoEveryone with “only” 64GB of RAM is pouting today, including me
- Taek 4y agoYou can run llama using 4 bits per parameter, 64 GB of RAM is more than enough
- geysersam 4y ago4 bits is ridiculously little. I'm very curious what makes these models so robust to quantization.
- MacsHeadroom 4y agoRead The Case for 4 Bit Precision. https://arxiv.org/abs/2212.09720 https://arxiv.org/abs/2212.09720 Spoiler: it's the parameter count. As parameter count goes up, but depth matters less. It just so happens that at around 10B+ parameters you can quantize down to 4bit with essentially no downsides. Models are that big now. So there's no need to waste RAM by having unnecessary precision for each parameter.
- Taek 4y agoFor completeness, there's also another paper that demonstrated you get more power/accuracy per-bit at 4 bits than at any other level of precision (including 2 bits and 3 bits)
- MacsHeadroom 4y agoThat's the paper I referenced. But newer research is already challenging it. 'Int-4 llama is not enough [0] - Int-3 and beyond' suggests 3-bit is best for models larger than ~10B parameters when combining binning and GPTQ. [0] https://nolanoorg.substack.com/p/int-4-llama-is-not-enough-int-3-and https://nolanoorg.substack.com/p/int-4-llama-is-not-enough-i...
- metadat 4y agoWhat if you have around 400GB of RAM? Would this be enough?
- gymbeaux 4y agoWhat I'm referring to requires around 67GB of RAM. With 400GB I would imagine you are in good shape for running most of these GPT-type models.
- numpad0 4y agoMore like finally "proven right" to have needlessly kept feeding 4/5th of 64GB to Chrome since 2018