5 ms·
I don't know the licensing and all that jazz (even if you self-host for your personal use it shouldn't matter). But, this paper[0] released a week ago claims "
by jonathan-adly 3y ago
I don't know the licensing and all that jazz (even if you self-host for your personal use it shouldn't matter). But, this paper[0] released a week ago claims " 99.3% of the performance level of ChatGPT while only requiring 24 hours of finetuning on a single GPU" (QLORA).
A quick test of the huggingface demo gives reasonable results[1]. The actual model behind the space is here[2], and should be self-hostable with reasonable effort.
0. https://arxiv.org/abs/2305.14314 https://arxiv.org/abs/2305.14314
1. https://huggingface.co/spaces/uwnlp/guanaco-playground-tgi https://huggingface.co/spaces/uwnlp/guanaco-playground-tgi
2. https://huggingface.co/timdettmers/guanaco-33b-merged https://huggingface.co/timdettmers/guanaco-33b-merged
- MacsHeadroom 3y agoYou link to Guanaco-33B, but Guanaco-65B is much more capable. CPU Version: https://huggingface.co/TheBloke/guanaco-65B-GGML https://huggingface.co/TheBloke/guanaco-65B-GGML GPU Version: https://huggingface.co/TheBloke/guanaco-65B-HF https://huggingface.co/TheBloke/guanaco-65B-HF 4bit GPU Version: https://huggingface.co/TheBloke/guanaco-65B-GPTQ https://huggingface.co/TheBloke/guanaco-65B-GPTQ
- wing-_-nuts 3y agoIt irritates me to no end that people don't list the system requirements of various models. How much ram and vram does one need to run 4,13,33,65B models at a reasonable speed? edit: instead I'll ask this, what's the best model to run on a system with a 24gb 4090 and 64gb of ram?
- orost 3y agoYou can just barely fit a 33B GPTQ model in 24GB VRAM. It will be in 4-bit mode, and without maximum context size, but it will be quite fast. Or you can run from RAM+VRAM in GGML format with llama.cpp (or a derivative), which will easily fit 65B models even at 5 or 8 bits, but at much lower speed.
- evanchisholm 3y agoIt's fairly simple to estimate ram requirements based on parameter count: https://blog.eleuther.ai/transformer-math/ https://blog.eleuther.ai/transformer-math/
- TheUninformed 3y agoA 4-bit quantized 33B parameter model will fit on your GPU and you'll be able to use a 2048 token context too. (4-bit quantized larger models are better than smaller 8bit/16bit models) You can run 4-bit quantized 65B models on your cpu, but it is slow, 1-2 tokens a second instead of 8-15 people typically get with a gpu, but you need two 24gb or an enterprise card with 48gb of ram to load them there. https://old.reddit.com/r/LocalLLaMA/wiki/models https://old.reddit.com/r/LocalLLaMA/wiki/models has the information you are irritated about not being listed.
- ac29 3y agoThe first link there gives very specific RAM requirements for each of the models. "Reasonable speed" is subjective, though.
- coolspot 3y agoGuanaco is indeed very capable and can replace GPT 3.5 in almost all scenarios, based on my tests. Easy way to self-host it is to use text-generation-webui[1] and 33B 4-bit quantized GGML model from TheBloke[2]. [1] https://github.com/oobabooga/text-generation-webui https://github.com/oobabooga/text-generation-webui [2] https://huggingface.co/TheBloke/guanaco-33B-GGML https://huggingface.co/TheBloke/guanaco-33B-GGML
- londons_explore 3y agoTable 1 in your link (https://arxiv.org/pdf/2305.14314.pdf https://arxiv.org/pdf/2305.14314.pdf) is a good way to compare models. I'd really like someone to make a big leaderboard/ranking engine which pits all these engines against eachother and publishes the resulting Elo score.
- jlund-molfese 3y agoLike https://lmsys.org/blog/2023-05-03-arena/ https://lmsys.org/blog/2023-05-03-arena/ ?
- the__prestige 3y agohttps://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
- _laiq 3y agoIncredible how people are underestimating ChatGPT or overestimating open-source models. A basic question: How can i join with SQL a column to a string separated with comma GPT-4: PostgreSQL: SELECT STRING_AGG(columnName, ', ') FROM tableName; Guanaco: Here is an example of how you could use the CONCAT function in MySQL to concatenate a string to a column value in a single-line query: SELECT CONCAT('The total price for ', product_name, ’ is ', SUM(price)) AS total FROM products; This will result in output like “The total price for Chocolate Bar is 10” or similar depending on your data. Yeah, no...
- yread 3y agoI also prefer the first but your point would be stronger if the example from Guanaco was actually wrong or internally inconsistent
- dustypotato 3y agoI have chatGPT with GPT-4 model. It never gives such concise answers. Is it because you used the API?
- laratied 3y agoThe default is rather wordy. You could instruct chatGPT4 to be less wordy or in the API put instructions to be less wordy in the system prompt or limit answers to a sentence. I would assume the poster just edited this though.
- seszett 3y agoMy test for open models is surprisingly simple. I just ask "What is the capital of France?" and I haven't had a correct answer yet in any model I tried. They often get Paris right at least, but most other details are wrong. Guanaco says: > The current capital of France is Paris. It has been so since 1982 when it replaced the previous one which was Vichy.
- simonw 3y agoI just got "The capital of France is Paris." from vicuna-v1-7B running entirely on my iPhone (using the MLC Chat app).
- deleted 3y ago[deleted]