2 ms·
> On throughput-focused benchmarks, Tokasaurus can outperform vLLM and SGLang by up to 3x+. Looks like they don't compare to TensorRT-LLM throughput numbers wh
by nabakin 1y ago
> On throughput-focused benchmarks, Tokasaurus can outperform vLLM and SGLang by up to 3x+.
Looks like they don't compare to TensorRT-LLM throughput numbers which, last I checked, are SOTA in open source.
- andersa 1y agoTensorRT-LLM being open source is a lie, all the important kernels are loaded from cubins.
- nabakin 1y agoYeah you're right (although, they started to open source some of that recently iirc). I meant SOTA for inference engines we can actually download and use ourselves.
- qeternity 1y agoIt also appears that this was a sampling benchmark...which is not representative. Generation benchmark was 5% faster than SGLang.