3 ms·
Nah, vram is definitely the main constraint for most people trying to do local inference of LLMs. If you look at a lot of the local LLM communities, for people
by Me1000 3y ago
Nah, vram is definitely the main constraint for most people trying to do local inference of LLMs. If you look at a lot of the local LLM communities, for people who aren't super interested in training, many people suggest the M2 Ultra or M3 Max with > 128GB of unified RAM just because it has so much memory. Again, you might not be able to train or fine tune as well, but the first step is just being able to keep the weights in memory. Inference isn't that computationally intensive relative to just how much RAM you need.
- brucethemoose2 3y agoYeah. I didn't mention them because the 64GB+ configs are very expensive, and I don't think you can even finetune on them.
- Me1000 3y agoThey're not cheap, but they're not _that_ expensive compared to buying four NVLink'd Nvidia cards with a combined similar amount of VRAM. Plus you get a whole computer with it and you don't have to worry about casings and power supplies, etc. But yeah, you will be more compute constrained if you go down that route.
- brucethemoose2 3y agoThe price is similar to 2x RTX 8000s, or even A6000s, but yeah your point stands. Power efficiency is something too. You run into immense pain the moment you venture outside of llama inference though.