3 ms·
We used an A100-80GB GPU. We didn't compare explicitly to Huggingface TGI but I think you should be able to compare the tokens/s achieved. One note is that thi
by chillee 3y ago
We used an A100-80GB GPU. We didn't compare explicitly to Huggingface TGI but I think you should be able to compare the tokens/s achieved.
One note is that this release is optimized for latency, while I think HF TGI might be more optimized for throughput.