5 ms·
Having installed this, this is an incredibly then wrapper around the following github repos: https://github.com/NVIDIA/trt-llm-rag-windows https://github.com/N
by operator-name 3y ago
Having installed this, this is an incredibly then wrapper around the following github repos:
https://github.com/NVIDIA/trt-llm-rag-windows https://github.com/NVIDIA/trt-llm-rag-windows
https://github.com/NVIDIA/TensorRT-LLM https://github.com/NVIDIA/TensorRT-LLM
It's quite a thin wrapper around putting both projects into %LocalAppData%, along with a miniconda environment with the correct dependnancies installed. Also for some reason the LLaMA 13b (24.5GB) and Ministral 7b (13.6GB) but only installed Ministral?
Ministral 7b runs about as accurate as I remeber, but responses are faster than I can read. This seems at the cost of context and variance/temperature - although it's a chat interface the implementation doesn't seem to take into account previous questions or answers. Asking it the same question also gives the same answer.
The RAG (llamaindex) is okay, but a little suspect. The installation comes with a default folder dataset, containing text files of nvidia marketing materials. When I tried asking questions about the files, it often cites the wrong file even if it gave the right answer.
- kkielhofner 3y agoThe wrapping of TensorRT-LLM alone is significant. I’ve been working with it for a while and it’s… Rough. That said it is extremely fast. With TensorRT-LLM and Triton Inference Server with conservation performance settings I get roughly 175 tokens/s on an RTX 4090 with Mistral-Instruct 7B. Following commits, PRs, etc I expect this to increase significantly in the future. I’m actually working on a project to better package Triton and TensorRT-LLM and make it “name and model and press enter” level usable with support for embeddings models, Whisper, etc.
- FirmwareBurner 3y ago>LLaMA 13b (24.5GB) and Ministral 7b (13.6GB) But the HW requirements state 8GB of VRAM. How do those models fit in that?
- av3csr 3y agoThey are int4 quantized
- Kranar 3y agoDoes int4 mean 4 bits per integer, or 4 bytes/32-bits. If it means that weights for an LLM can be 4 bits well that's just mind boggling.
- sillysaurusx 3y agoFour bits per parameter. (A parameter is what you call an integer here.) I was skeptical of it for some time, but it seems to work because individual parameters don’t encode much information. The knowledge is embedded thanks to having a massive number of low bit parameters.
- justahuman74 3y ago4 bits