2 ms·
It's not so much about preference but controlling our load and resource consumption right now. We're setting an easy threshold to meet consistently and the adde
by TrueDuality 3y ago
It's not so much about preference but controlling our load and resource consumption right now. We're setting an easy threshold to meet consistently and the added delay allows us to imperceptibly handle things like crashes in Nvidia's drivers, live swapping of model and LoRA layers, etc.
(For clarification the users preference in my original post, is about interactive users preferring to see a stream of tokens coming in rather than waiting for the entire request to complete and having it show up all at once. The performance of that sets the expectation for the time of non-interactive responses.)