3 ms·
7 tokens per sec on an i5-11400 CPU using llama.cpp - that's pretty real time for personal use I would think.
by programd 3y ago
7 tokens per sec on an i5-11400 CPU using llama.cpp - that's pretty real time for personal use I would think.
3 ms·