3 ms·
But how fast? I see other companies advertising 1.5 second response times for GPT-J, but a now assume that’s average per token, as for, say, a 200 word prompt r
by d13 5y ago
But how fast? I see other companies advertising 1.5 second response times for GPT-J, but a now assume that’s average per token, as for, say, a 200 word prompt response times can be well over a minute during heavy use periods like weekends when everyone is hitting their side projects.
- varunkmohan 5y agoFor the public API, we should be under 100ms per token but don't have a strict guarantee. If you have a strict SLA, you can talk to us and we can get it to as low as 20ms per token at high load.