3 ms·
Yeah, This can definatly be used for local models, but the problem is that most personal computers cannot host large LLMs and the cost is not cheaper than close
by joeyxiong 3y ago
Yeah, This can definatly be used for local models, but the problem is that most personal computers cannot host large LLMs and the cost is not cheaper than closed LLMs. But for organisations, local LLMs are a better choice.
- stevenhuang 3y agoIt's a lot closer these days with 30B 4bit quantized GPTQ fitting in 1 RTX 3090. Goes from 20 tokens per second to 15 tokens per second nearing the ~3k token context length, with similar quality in output to chatgpt 3.5.
- bytefactory 3y agoI think local LLMs are great for tinkerers, and with quantization can run on most modern PCs. I am not comfortable sending over my personal data over to OpenAI/Anthropic, so I've been playing around with https://github.com/PromtEngineer/localGPT/ https://github.com/PromtEngineer/localGPT/, GPT4All, etc. which keep the data all local. Sliding window chunking, RAG, etc. seem more sophisticated than the other document LLM tools, so I would love to try this out if you ever add the ability to run LLMs locally!