4 ms·
Two things I'm curious to know: 1. How many tokens can 'traditional' models (e.g. Mistral's 8x7B) fit on a single 80GB GPU? 2. How does quantization affect the
by nostrowski 3y ago
Two things I'm curious to know:
1. How many tokens can 'traditional' models (e.g. Mistral's 8x7B) fit on a single 80GB GPU?
2. How does quantization affect the single transformer layer in the stack? What are the performance/accuracy trade-offs that happen when so little of the stack depends on this bottleneck?
- patrakov 3y agoMixtral 8x7b runs well (i.e., produces the correct output faster than I can read it) on a modern AMD or Intel laptop without any use of a GPU - provided that you have enough RAM and CPU cores. 32 GB of RAM and 16 hyperthreads are enough with 4-bit quantization if you don't ask too much in terms of context. P.S. Dell Inspiron 7415 upgraded to 64 GB of RAM here.