3 ms·
While I’m very excited for an LLM powered Siri that has access to my calendar, email, and messages, it’s worth noting that llama.cpp is using metal for accelera
by ericskiff 3y ago
While I’m very excited for an LLM powered Siri that has access to my calendar, email, and messages, it’s worth noting that llama.cpp is using metal for acceleration and the M2 Mac Studio is one of the best values out there for running very large models like llama2 70B derivatives or Falcon 180B. The shared ram/vram allows you to load huge models and it’s got decent tokens/sec performance.
In the last few weeks, the initial prompt evaluation step got much faster with llama.cpp so everything that depends on it like the python library and lmstudio are faster as well. I’m very happy with my purchase at this point and it seems to keep getting better.
- verdverm 3y agoI'm curious about this memory things with the Macs. I was looking at Airs with M2, and the memory is only 8G. Is that the limit, if I'm say... running containers and applications? Where can I learn more? I don't understand how one can ship a computer today with only 8G of mem. My phone has more than that. Are developers able to work with this amount of mem, or are they forking out $3k + for a Pro with 24G? I went with a Framework because I got good hardware for way less money
- lbourdages 3y agoMy previous MBP had 32 GB of ram and it was constantly swapping. Now I have 64, and it's holding up. I guess it really depends on your work environment, but I need Docker, IntelliJ, plus a bunch of other stuff, and it quickly adds up. I don't know how anyone can use a laptop with only 8GB of ram. Maybe if all they do on it is Facebook and Youtube ¯\_(ツ)_/¯.