3 ms·
You already can run a local llama instance on a high-end graphics card (6+ GB VRAM).
by nsvd 3y ago
You already can run a local llama instance on a high-end graphics card (6+ GB VRAM).
- tpmx 3y agoAnd it's hilariously bad (in comparison to regular chatgpt).
- Der_Einzige 3y agoAnd slow. They never tell you that quantization of many LLMs slows down your inference, sometimes by orders of magnitude.
- arugulum 3y agoIt depends on the quantization method, but yes some of the most commonly used ones are extremely slow.
- nvy 3y agoYes, I can, but (see my edit) there's very little utility because the quality of output is very low. Frankly anything worse than the ChatGPT-3.5 that runs on the "open"AI free demo isn't much of a tool.