4 ms·
Is LMDeploy the Ultimate Solution? Why It Outshines VLLM, TRT-LLM, TGI, and MLC
- timliu9 2y agoWhy was onnx not part of the tested runtimes? Seems like an oversight
- chaoyu 2y agoonnx is not a good option for LLM type of autoregressive generation
- helloericsf 2y agoPersonally, I never seen onnx used for LLM.
- ssheng 2y agoHow does Exllama rank among these? Heard good things about it.
- helloericsf 2y agoSeems interesting! https://github.com/turboderp/exllama https://github.com/turboderp/exllama "A more memory-efficient rewrite of the HF transformers implementation of Llama for use with quantized weights."
- helloericsf 2y ago4-bit quantization tends to come at the cost of output quality losses. https://github.com/ggerganov/llama.cpp/issues/9 https://github.com/ggerganov/llama.cpp/issues/9
- ssheng 2y agoQuality loss with quantization is expected. It seems like with GPTQ the loss is within acceptable range based on the perplexity score shown.
- ShawnBasquiat 2y agoWhy aren't there more of these benchmark studies? How did TGI make the cut?
- deleted 2y ago[deleted]