4 ms·
The linked paper also includes performance of a second PC configuration. The 'high' configuration is a 4090 (24GB) with a 24-core cpu (iirc), 192 GB ram, pci 4.
by tofof 3y ago
The linked paper also includes performance of a second PC configuration. The 'high' configuration is a 4090 (24GB) with a 24-core cpu (iirc), 192 GB ram, pci 4. The 'low' configuration is a 2080ti (11GB) with an 8-core cpu, 64 GB ram, pci 3. The low configuration still averaged a 5x speedup over llama.cpp, so yeah, this is exciting. They noted that for particularly large models (60B+) the hot neurons don't all fit into the gpu's vram and performance falls off. Similarly, for small models with small context sizes enough fits into the gpu to begin with that the performance gains are less pronounced. So there's going to be a sweet spot with regard to the combination of model and context size for a particular configuration, but yes, this still gives a huge speedup compared to llama.cpp.
Image of the PC-low configuration's results: https://i.imgur.com/X5reGkd.png https://i.imgur.com/X5reGkd.png showing speedups ranging from 2x-13x.
- LoganDark 3y ago> 24-core cpu Pretty sure it's an 8-core CPU, or at least that's how they're using it. Their demo videos show only 8 threads in use, probably because the "efficiency cores" would negate some of the performance win.