3 ms·
I'm not entirely sure your statements are correct in today's landscape. I run models both on a M1 chip and 3090. The token generation are about the same if I us
by syntaxing 3y ago
I'm not entirely sure your statements are correct in today's landscape. I run models both on a M1 chip and 3090. The token generation are about the same if I use huggingface transformers on the 3090 or llama.cpp on the M1 with the same quantization (4 bit, I don't have enough RAM to run 8 bit on the M1). Most major local LLM libraries support Metal. I know you can use MPS backend on huggingface transformers but I have not tried it since llama.cpp (Ollama to be more specific) works great on a M1 to begin with.
- smoldesu 3y ago> are about the same if I use huggingface transformers on the 3090 or llama.cpp on the M1 with the same quantization What is the difference when you use llama.cpp or transformers for both? In my experience, the difference is pretty much far-and-away once you get the model loaded into RAM. My 3070 inferences about as fast as I've seen the M1 Ultra go (and my system is using a $500 chip, not a $5,000 one). For the money, I think you're just not getting the most bang for your buck with Apple Silicon. It's fine if you want a Mac and use AI to justify the bigger one, but... it would be such a sore experience to spend $7,000 to avoid CUDA. It's the wrong walled-garden to pick if you care more about AI. > (4 bit, I don't have enough RAM to run 8 bit on the M1). That does make a big difference in inferencing speed, though. Lower precision models are much faster, if you ran 4bit on both machines it would probably close the gap significantly.