11 ms·
State-of-the-Art Chatbot, Vicuna-7B, now runs on MacBook with GPU acceleration
- zhisbug 3y agoNo llama.cpp nor any compilation complexity. Run with two Python commands!
- superkuh 3y agoI think you have it backwards. The python (ie, huggingface, etc) implementations of transformers are the complex ones with dependency hell so bad even there's even a layer of package manager / env hell. This version of fastchat (there's 2) required a particular commit of huggingface libs for quite a while. Something that only changed recently. And it'll happen again in the future. Python just hides this complexity... until it doesn't. Like beautiful but rapidly rotting fruit. llama.cpp will remain a single two line project (git clone https://github.com/ggerganov/llama.cpp https://github.com/ggerganov/llama.cpp, make -j) that will compile easily and run on anything. No external deps to pin to a particular commit (that will only have a lifetime of some months) as things change rapidly. That said, the changes in the ggml weights format the last 2 weeks were annoying, but now that the mmap-style weights are settled on it should be less converting. In that sense huggingface wins, it only has two incompatible weights formats. llama.cpp's ggml has had 3.
- Casteil 3y agoThis has been my experience so far as well. GPT4All feels pretty fragile with all its dependencies.
- sterlind 3y agoI've spent the past couple days packaging an LLM playground environment as a Nix expression. it's been pure hell. also nice to see you again, superkuh. I frequented your IRC channel about a decade ago.
- siraben 3y agoHave you been successful in getting the LLM playground up with Nix?
- sterlind 3y agoyes, almost! I used poetry2nix and grafted a bunch of overrides to fix the torch-2.0 build, and I just got cuda working with it. I'm testing triton now. I'll submit my PR to poetry2nix so watch that space if you want it.
- superkuh 3y agoUsing nix and then complaining about having to set up your compilation environment libs/etc is kind of like sticking a rod in your bike's wheel spokes and complaining about crashing. Don't give up on the idea of system libraries (ie, use nix) and this doesn't happen. Also, hi? I don't recall you by that nick but the internet is a small place sometimes.
- sterlind 3y agooh, I'm very aware that I've brought this upon myself, but I'm sticking out for the greater good (and stubbornness.) specifically, I'm trying to benchmark a bunch of different GPU configurations on different workloads on vast.ai, which uses Docker containers. I abhor Dockerfiles and my experience building containers with nix has been pleasant, so that's what I'm doing and why. fortunately I think I'm getting past the learning curve. did our channel survive the demise of freenode? I was andares, I think I used to be annoying but I've gotten better.
- superkuh 3y agoAh. Hi! Yes. We still exist in the same place but on libera now.
- anotherhue 3y agoCare to share some of your progress? I have similar (stronger?) feelings regarding Dockerfile's big-ball-of-state nonsense. (The irony of holding this opinion while dealing with pre-trained AI models is not lost)
- zhisbug 3y agono, the requirement on a particular HF commit has been fixed. It is no longer needed.
- zhisbug 3y agoit is really a matter of having faith on pytorch (or JAX) or on third-party cross-platform supports like llama-cpp. Apparently pytorch reduces a lot of complexity and grows extremely faster on cross-platform supports. And, PyTorch does so well on GPUs!
- superkuh 3y agoRight. That particular problem has been fixed. But the fact that it was needed indicates it will happen again. It exposes the underlying complexity of the huggingface transformer stack. It's wonderful code, don't get me wrong. It's just the furthest thing possible from the least complex.
- sottol 3y agoDid they release the merged weights, yet? I'd love to try this model. Afaict from the docs, you still need to request the original Llama weights from Meta (or get ahold of them another way), then apply the diff-weights requiring 60GB RAM?
- valine 3y agoYou can find the merged weights pretty easily online, I wouldn't hold your breath waiting for an official release given the licensing issues around LLaMA.
- youssefabdelm 3y agoSo far I think what these models lack is memory of people and other things. Especially if not as popular. And probably a ton more. E.g. try asking it "Who is Tyler Volk?" Then try asking GPT-4 "Who is Tyler Volk?" Then check who he is online.
- circuit10 3y agoProbably because the parameter count is way lower so it's less able to memorize things
- psychphysic 3y agoThe language is also quite unnatural feeling. Neat none the less but hardly a standout in my opinion. Everything is state of the art at the moment I guess so can't criticise that too much.
- Casteil 3y agoAnyone here who's used both this and GPT4All? Any thoughts/input on how they compare?
- weichiang 3y agoI asked GPT4All one of Vicuna's benchmark questions: "What if the Internet had been invented during the Renaissance period?" Check out their responses: https://imgur.com/a/mPrdZ1W https://imgur.com/a/mPrdZ1W More questions here: https://vicuna.lmsys.org/eval/ https://vicuna.lmsys.org/eval/ Note: not an apple-to-apple comparison but that's the model checkpoint I found on their git repo.
- superkuh 3y agoMy one take away after playing with both chat mode and text completion modes is that gpt4all 7B 4bit stays on the chat rails (doesn't start taking the role of the user, or spewing fine tuning boilerplate) much better than vicuna 7B 4bit. In text completion they're about the same but I'd still prefer the vanilla llama 7B in that case. There are a couple versions of gpt4all fine-tuned llama 7B and my favorite is the unfiltered one (gpt4all-lora-unfiltered-quantized.bin). https://github.com/nomic-ai/gpt4all#try-it-yourself https://github.com/nomic-ai/gpt4all#try-it-yourself
- zhisbug 3y agoLmsys hasn't released any official 4-bit version. It might be a better idea to wait for the official 4-bit version. But it is interesting to learn that the third-party 4bit version has performance degeneration.
- superkuh 3y agoLmsys hasn't released any official weights for anything. They've released "deltas" and other people have applied those deltas to the appropriate llama weights and done the quantization. I reject your premise that the 8 to 4 bit quantization is the cause of the vicuna fine-tuned llamas very average performance though. This hasn't been the case for any of the other 8 to 4 bit quantizations. It would be a unique outlier. And so I don't think this is the "cause" here.
- wejick 3y agoSeems like llama derived model are flourishing. However with llama is licensed as academic only and noncommercial model, what is the path for bringing this to production of for profit purpose? I certainly interested doing so.
- alwayslikethis 3y agoCopilot style. Train a distilled model based on it, and now it's a new model unencumbered by copyright.
- ReptileMan 3y agoThe Silicon Valley ethos has always been - do it first worry about legality later. If you go bust - nobody will care. If you become small - you will be ignored. If you go big - lawyers will figure something out to cut a deal.
- JohnFen 3y agoThat is a thoroughly bankrupt ethos that should be denounced every time it pops up. It is literally condoning criminality.
- deleted 3y ago[deleted]
- fortyseven 3y ago"Won't somebody think of the poor defenseless corporations?!"
- baq 3y agoNo crimes in this case, just license breaches. After a few training iterations, it’ll be very muddled anyways.
- JohnFen 3y agoYes, I was speaking of the general ethos, not a specific case. But let's take Uber as an example of that ethos in action -- Uber committed actual crimes as part of their growth strategy.
- tric 3y agoWhy is there so much focus on running GPT models on Mac OS? Is there something special about Apple's new chip, or Mac OS?
- 19h 3y agoUnified memory allows both CPU and GPU to use the same memory, effectively giving a MacBook with 96GB of memory 96GB of VRAM (minus OS overhead obv).
- matwood 3y agoThe shared ram and neural engine make for an interesting/powerful platform if people are willing to port to it.
- steve_adams_86 3y agoAre the neural engines able to be leveraged by 3rd parties yet? I thought there was no API available yet.
- rnosov 3y agoThey are leveraging Apple’s Metal Performance Shaders[1] not the neural engine. From the chart, it looks like you might get ~20x max boost on inference over plain CPU. Obviously, it's not like having RTX 4090 but better than nothing. [1] https://pytorch.org/blog/introducing-accelerated-pytorch-training-on-mac/ https://pytorch.org/blog/introducing-accelerated-pytorch-tra...
- wmf 3y agoCoreML is the API.
- wmf 3y agoApple's unified memory should allow running large models like 65B that will not fit on a consumer GPU, but mostly I see people talking about the smaller 7B sizes that can run anywhere.
- 3y ago
- bawana 3y agoMacBook with M1 chip here.python installed with homebrew tried to install with: pip install fschat then tried to run it with: python3 -m fastchat.serve.cli --model -name vicuna-7b --device mps --load-8bit got this: traceback (most recent call last): File "<frozen runpy>", line 198, in _run_module_as_main File "<frozen runpy>", line 88, in _run_code File "/opt/homebrew/lib/python3.11/site-packages/fastchat/serve/cli.py", line 9, in <module> from transformers import AutoTokenizer, AutoModelForCausalLM, LlamaTokenizer ModuleNotFoundError: No module named 'transformers' so I did this: pip install transformers command tried again: python3 -m fastchat.serve.cli --model -name vicuna-7b --device mps --load-8bit got: Traceback (most recent call last): File "/opt/homebrew/lib/python3.11/site-packages/transformers/utils/import_utils.py", line 1126, in _get_module return importlib.import_module("." + module_name, self.__name__) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/opt/homebrew/Cellar/python@3.11/3.11.2_1/Frameworks/Python.framework/Versions/3.11/lib/python3.11/importlib/__init__.py", line 126, in import_module return _bootstrap._gcd_import(name[level:], package, level) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "<frozen importlib._bootstrap>", line 1206, in _gcd_import File "<frozen importlib._bootstrap>", line 1178, in _find_and_load File "<frozen importlib._bootstrap>", line 1128, in _find_and_load_unlocked File "<frozen importlib._bootstrap>", line 241, in _call_with_frames_removed File "<frozen importlib._bootstrap>", line 1206, in _gcd_import File "<frozen importlib._bootstrap>", line 1178, in _find_and_load File "<frozen importlib._bootstrap>", line 1149, in _find_and_load_unlocked File "<frozen importlib._bootstrap>", line 690, in _load_unlocked File "<frozen importlib._bootstrap_external>", line 940, in exec_module File "<frozen importlib._bootstrap>", line 241, in _call_with_frames_removed File "/opt/homebrew/lib/python3.11/site-packages/transformers/models/__init__.py", line 15, in <module> from . import ( File "/opt/homebrew/lib/python3.11/site-packages/transformers/models/mt5/__init__.py", line 29, in <module> from ..t5.tokenization_t5 import T5Tokenizer File "/opt/homebrew/lib/python3.11/site-packages/transformers/models/t5/tokenization_t5.py", line 26, in <module> from ...tokenization_utils import PreTrainedTokenizer File "/opt/homebrew/lib/python3.11/site-packages/transformers/tokenization_utils.py", line 26, in <module> from .tokenization_utils_base import ( File "/opt/homebrew/lib/python3.11/site-packages/transformers/tokenization_utils_base.py", line 74, in <module> from tokenizers import AddedToken File "/opt/homebrew/lib/python3.11/site-packages/tokenizers/__init__.py", line 80, in <module> from .tokenizers import ( ImportError: dlopen(/opt/homebrew/lib/python3.11/site-packages/tokenizers/tokenizers.cpython-311-darwin.so, 2): no suitable image found. Did find: /opt/homebrew/lib/python3.11/site-packages/tokenizers/tokenizers.cpython-311-darwin.so: mach-o, but wrong architecture /opt/homebrew/lib/python3.11/site-packages/tokenizers/tokenizers.cpython-311-darwin.so: mach-o, but wrong architecture The above exception was the direct cause of the following exception: Traceback (most recent call last): File "<frozen runpy>", line 198, in _run_module_as_main File "<frozen runpy>", line 88, in _run_code File "/opt/homebrew/lib/python3.11/site-packages/fastchat/serve/cli.py", line 9, in <module> from transformers import AutoTokenizer, AutoModelForCausalLM, LlamaTokenizer File "<frozen importlib._bootstrap>", line 1231, in _handle_fromlist File "/opt/homebrew/lib/python3.11/site-packages/transformers/utils/import_utils.py", line 1116, in __getattr__ module = self._get_module(self._class_to_module[name]) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/opt/homebrew/lib/python3.11/site-packages/transformers/utils/import_utils.py", line 1128, in _get_module raise RuntimeError( RuntimeError: Failed to import transformers.models.auto because of the following error (look up to see its traceback): dlopen(/opt/homebrew/lib/python3.11/site-packages/tokenizers/tokenizers.cpython-311-darwin.so, 2): no suitable image found. Did find: /opt/homebrew/lib/python3.11/site-packages/tokenizers/tokenizers.cpython-311-darwin.so: mach-o, but wrong architecture /opt/homebrew/lib/python3.11/site-packages/tokenizers/tokenizers.cpython-311-darwin.so: mach-o, but wrong architecture
- dchuk 3y agoSo if I have a 32GB RAM Macbook Pro, and the instructions say this: "Vicuna-13B This conversion command needs around 60 GB of CPU RAM." Does this mean I simply cannot run that model at all? Or will it rip into HD swap or something to make the model weights and just take forever?
- UncleOxidant 3y agoI just reached that step on my Linux laptop which has 32GB of RAM. I'm about to give it a try anyway, but I'm not hopeful based on that comment. I'm wondering if anyone is torrenting these Vicuna-13B weights?
- MMMercy2 3y agoYou can try the smaller 7B version.
- acchow 3y agoCan someone explain why computing a delta needs to hold the entire model at once? Can't it just do one layer at time?
- GaggiX 3y agoSomeone really needs to write a script that does not load both entire models into memory to do this.
- FLT8 3y agoVicuna-13B loads and idles at ~26GB RAM usage on a M1Max/64GB. When answering questions, that grows to around 75GB, and yes, you can feel it (and the machine) slow down significantly when it starts hitting swap. I think realistically you'd be wanting to stick to the 7B model on a 32G machine (even if you could get the weight deltas to apply correctly).
- syntaxing 3y agoI’ve been using the GPTQ 4 bit quantized 13B with Text generation web UI and it’s been amazing. Probably the closest to ChatGPT I have used so far. I still get an issue where it keeps on talking to itself by generating its own prompt and then answering it. Has anyone experienced the same thing?
- brookst 3y agoI have not experienced that problem, but it sounds both annoying and funny. How often do you encounter it? A few times a day but it varies quite a bit. Have you been able to tell what causes it? Shorter prompts sometimes cause it but I have seen it on longer prompts as well. Is there a new version that fixes the issue? Not that I've seen released but it is worth checking.
- syntaxing 3y agoIt happens randomly, and I tried adjusting the gradio+model settings to match FastChat. I should start taking some screenshots of these cause some are funny. I asked how do I update a git repo and it answered correctly with git pull. Then it added HUMAN: how do I delete the whole folder and start over and answered ASSISTANT: try using git reset —hard, if not use rm -rf (paraphrasing here).
- Cyphase 3y agoThat was brilliant. Thanks! You're welcome.
- UncleOxidant 3y agoI've gotten that which alpaca.cpp. It'll start asking itself questions and then answer them.
- tyfon 3y agoI've been testing quite a few of these models lately. For me, the absolute best is still the 65B 4-bit quantized llama model with the correct prompt and parameters, both for programming, language and general questions. I am actually getting about 2 tokens/second with the latest llama.cpp using 16 threads on a 5950x with 64 gb of ram. 16 threads seems to be the sweet spot, any higher and it slows down, any lower and it is less consistent in the time to produce a token. I am 100% convinced that the AI "market" will be a local thing. Running this and having access to all the information stored in the weights easily and without internet is just so great I think :) Edit: the responding to itself "bug" is most likely an issue with the prompt you issue. The recent llama.cpp has a good starting point in the examples/chat-13b.sh I am using a modified version of that where I set the 65B model, change the moscow stuff to cairo and the node.js to a small C program.
- ThorsBane 3y agoThis is cool, but how can we get Facebook’s fingers out of the pie with open source weights?
- deleted 3y ago[deleted]