3 ms·Faster Mixtral inference with TensorRT-LLM and quantization2 points by tikkun 3y agotikkun 3y agoTensorRT-LLM plus int8 quantization = cheaper, minimal quality loss. Some nice tables in the post