4 ms·
For people who saw this and might want a recomendation, I like running a tiny qwen model with llama cpp. Qwen2.5 coder 0.5B or 1.5B (not the instruct version)
by jboss10 3mo ago
For people who saw this and might want a recomendation, I like running a tiny qwen model with llama cpp. Qwen2.5 coder 0.5B or 1.5B (not the instruct version)
On a modern-ish GPU these should run really fast with little latency. They cost nothing and don't send your data to anyone.