3 ms·
I’m running mistral 7B on a M1 Mac 8GB just barely. It’s ask a question get a coffee type of thing. No idea how this works, as 32 bit floats require 4 bytes and
by RandomWorker 3y ago
I’m running mistral 7B on a M1 Mac 8GB just barely. It’s ask a question get a coffee type of thing. No idea how this works, as 32 bit floats require 4 bytes and with 7B it would need to be swapping with the SSD.
If I had the cash I would go for 24GB M2/3 pro. That would allow me to comfortably load the 7B model in to ram.
- puchatek 3y agoCan I ask what you're using it for?
- RandomWorker 3y agoNo, that’s the whole point of running this stand alone without anyone spying.
- m1sta_ 3y agoI run mistral on an M2 air and it's broadly similar to chatgpt.
- wenc 3y agoHow? I have an M2 Pro and I run 7B and 13B models through Ollama and also LM Studio. Because there’s no CUDA, the speed is much slower than ChatGPT. The answers from 7B are also not at the same quality as ChatGPT. (Lots of mistakes and hallucinations)
- evnc 3y agoHave you looked into quantization? At 8-bit quantization, a 7B model requires ~7GB of RAM (plus a bit of overhead); at 4-bit, it would require around 3.5GB and fit entirely into the RAM you have. Quality of generation does degrade a bit the smaller you quantize, but not as much as you may think.
- RandomWorker 3y agoThis is interesting; I've written how I set it up here; https://christiaanse.ca/posts/running_llm/ https://christiaanse.ca/posts/running_llm/