4 ms·
This looks quite cool! It's basically a tech demo for TensorRT-LLM, a framework that amongst other things optimises inference time for LLMs on Nvidia cards. The
by operator-name 3y ago
This looks quite cool! It's basically a tech demo for TensorRT-LLM, a framework that amongst other things optimises inference time for LLMs on Nvidia cards. Their base repo supports quite a few models.
Previously there was TensorRT for Stable Diffusion[1], which provided pretty drastic performance improvements[2] at the cost of customisation. I don't forsee this being as big of a problem with LLMs as they are used "as is" and augmented with RAG or prompting techniques.
[1]: https://github.com/NVIDIA/Stable-Diffusion-WebUI-TensorRT https://github.com/NVIDIA/Stable-Diffusion-WebUI-TensorRT
[2]: https://reddit.com/r/StableDiffusion/comments/17bj6ol/hows_your_results_with_tensorrt/ https://reddit.com/r/StableDiffusion/comments/17bj6ol/hows_y...
- operator-name 3y agoHaving installed this, this is an incredibly then wrapper around the following github repos: https://github.com/NVIDIA/trt-llm-rag-windows https://github.com/NVIDIA/trt-llm-rag-windows https://github.com/NVIDIA/TensorRT-LLM https://github.com/NVIDIA/TensorRT-LLM It's quite a thin wrapper around putting both projects into %LocalAppData%, along with a miniconda environment with the correct dependnancies installed. Also for some reason the LLaMA 13b (24.5GB) and Ministral 7b (13.6GB) but only installed Ministral? Ministral 7b runs about as accurate as I remeber, but responses are faster than I can read. This seems at the cost of context and variance/temperature - although it's a chat interface the implementation doesn't seem to take into account previous questions or answers. Asking it the same question also gives the same answer. The RAG (llamaindex) is okay, but a little suspect. The installation comes with a default folder dataset, containing text files of nvidia marketing materials. When I tried asking questions about the files, it often cites the wrong file even if it gave the right answer.
- kkielhofner 3y agoThe wrapping of TensorRT-LLM alone is significant. I’ve been working with it for a while and it’s… Rough. That said it is extremely fast. With TensorRT-LLM and Triton Inference Server with conservation performance settings I get roughly 175 tokens/s on an RTX 4090 with Mistral-Instruct 7B. Following commits, PRs, etc I expect this to increase significantly in the future. I’m actually working on a project to better package Triton and TensorRT-LLM and make it “name and model and press enter” level usable with support for embeddings models, Whisper, etc.
- FirmwareBurner 3y ago>LLaMA 13b (24.5GB) and Ministral 7b (13.6GB) But the HW requirements state 8GB of VRAM. How do those models fit in that?
- av3csr 3y agoThey are int4 quantized
- Kranar 3y agoDoes int4 mean 4 bits per integer, or 4 bytes/32-bits. If it means that weights for an LLM can be 4 bits well that's just mind boggling.
- sillysaurusx 3y agoFour bits per parameter. (A parameter is what you call an integer here.) I was skeptical of it for some time, but it seems to work because individual parameters don’t encode much information. The knowledge is embedded thanks to having a massive number of low bit parameters.
- justahuman74 3y ago4 bits