3 ms·
https://nvidia.github.io/TensorRT-LLM/performance.html https://nvidia.github.io/TensorRT-LLM/performance.html It was one of the fastest backends last time I ch
by fisf 3y ago
https://nvidia.github.io/TensorRT-LLM/performance.html https://nvidia.github.io/TensorRT-LLM/performance.html
It was one of the fastest backends last time I checked (with vLLM and lmdeploy being comparable), but the space moves fast. It uses cuda under the hood, torch is not relevant in this context.