3 ms·
I run llama chat 70b on a p3 8x large (4 Tesla) and it runs at like 1-5 tokens per sec. And I’m running the model with only 4 bit precision. Are you doing anyth
by 2099miles 3y ago
I run llama chat 70b on a p3 8x large (4 Tesla) and it runs at like 1-5 tokens per sec. And I’m running the model with only 4 bit precision. Are you doing anything else?