4 ms·
What kind of latency do you observe when using Gemini with such large contexts?
by theolivenbaum 2y ago
What kind of latency do you observe when using Gemini with such large contexts?
- dmazin 2y agoInference takes 30+ seconds before it starts outputting even the first token.
- danielbln 2y agoHow's recall over 1M tokens?
- Workaccount2 2y agoAround 250k tokens its around 15-20 seconds. When you really pack it full (which I have only done just to test it) it can take a minute before it outputs anything. It also sometimes glitches out with massive contexts, and just stops generation mid sentence. Asking it to repeat the response clears it up though.