5 ms·
Yeah, it's CPU only, and it is using about 38g for 70B and 7g 7B. Guessing that is mostly from the large caches that it keeps around for efficiency. If you want
by srush 3y ago
Yeah, it's CPU only, and it is using about 38g for 70B and 7g 7B. Guessing that is mostly from the large caches that it keeps around for efficiency. If you wanted to pay some computational cost, you could likely get that down by quantizing activations. Llama2 has some of these tricks built in automatically, for instance grouped query attention.