3 ms·
Oh, I see, I misread. Thought it meant 0.6ms per output token. Now I get that it’s saying “prompt token”, so if your prompt is 100 tokens, that’s 60ms. That se
by elicksaur 2y ago
Oh, I see, I misread. Thought it meant 0.6ms per output token. Now I get that it’s saying “prompt token”, so if your prompt is 100 tokens, that’s 60ms.
That seems pretty fast. 1.6k tokens for a 1s time. Do other models compare to that? I’m not sure what the current top ranking for this metric looks like.
- lostmsu 2y agoLatency here is weird. You are using the number as bandwidth in this calculation. Perhaps reporter doesn't really know what he's talking about.