3 ms·
Yeah it's certainly possible, but it's not the focus of this implementation, which is more latency focused (so BS=1).
by chillee 3y ago
Yeah it's certainly possible, but it's not the focus of this implementation, which is more latency focused (so BS=1).
- lmeyerov 3y agoyeah i'm curious how this would stack up to vllm in a batch setting long-term, bodes well as both methodologies should be combineable, just curious for current point-in-time wrt being relevant for production use
- chillee 3y agoI wouldn't recommend using it for a batch serving setting today. One crucial optimization for batched serving (which you need if you have a large number of requests) is continual batching, which this implementation doesn't have.