3 ms·TensorRT-LLM plus int8 quantization = cheaper, minimal quality loss. Some nice tables in the postby tikkun 3y agoTensorRT-LLM plus int8 quantization = cheaper, minimal quality loss. Some nice tables in the post