3 ms·
If you stream the answer, the first token time is roughly the per token time.
by elicksaur 2y ago
If you stream the answer, the first token time is roughly the per token time.
- lostmsu 2y agoNo, you have to feed the entire prompt token-by-token before getting response.
- elicksaur 2y agoOh, I see, I misread. Thought it meant 0.6ms per output token. Now I get that it’s saying “prompt token”, so if your prompt is 100 tokens, that’s 60ms. That seems pretty fast. 1.6k tokens for a 1s time. Do other models compare to that? I’m not sure what the current top ranking for this metric looks like.
- lostmsu 2y agoLatency here is weird. You are using the number as bandwidth in this calculation. Perhaps reporter doesn't really know what he's talking about.