3 ms·
Nowhere near as neat as candle or ggml, but just released a 4-bit rust llama2 implementation with simd. Runs pretty fast. https://github.com/srush/llama2.rs/ h
by srush 3y ago
Nowhere near as neat as candle or ggml, but just released a 4-bit rust llama2 implementation with simd. Runs pretty fast.
https://github.com/srush/llama2.rs/ https://github.com/srush/llama2.rs/
- 1ba9115454 3y agoThat's really cool.
- unshavedyak 3y agoThat is _really_ cool. Would be interested in some requirements being posted. Eg RAM required for the 70B model if CPU, VRAM required if GPU, etc. edit: I know it's memory mapping, but that still loads data in RAM to execute so if you had 512MiB of RAM it would likely be slow as hell.
- srush 3y agoYeah, it's CPU only, and it is using about 38g for 70B and 7g 7B. Guessing that is mostly from the large caches that it keeps around for efficiency. If you wanted to pay some computational cost, you could likely get that down by quantizing activations. Llama2 has some of these tricks built in automatically, for instance grouped query attention.