3 ms·
tokens per second is normally the biggest cause of latency, so amortized I would bet a decent LLM on the phone is still slower than groq/aopenai etc
by nimchimpsky 2y ago
tokens per second is normally the biggest cause of latency, so amortized I would bet a decent LLM on the phone is still slower than groq/aopenai etc