4 ms·
VRAM is not the main constraint, is it? The computational power of any of the new graphical cards (beside from the highest end models, where the VRAM is the act
by treesciencebot 3y ago
VRAM is not the main constraint, is it? The computational power of any of the new graphical cards (beside from the highest end models, where the VRAM is the actual constraint, like RTX 4090s) is absurdly low on the stuff that actually matters (tensor cores, cuda cores, etc.). They are graphical cards, equipped with consumer grade VRAM (instead of HBM) and IMHO it will take a very big shift before we see them being used as real AI accelerators.
- sp332 3y agoWell if you don't have enough VRAM to hold the whole model, you have to swap out to main RAM for every iteration, which for most home/local setups means once for every single token generated. So having enough is kind of a minimum.
- Me1000 3y agoNah, vram is definitely the main constraint for most people trying to do local inference of LLMs. If you look at a lot of the local LLM communities, for people who aren't super interested in training, many people suggest the M2 Ultra or M3 Max with > 128GB of unified RAM just because it has so much memory. Again, you might not be able to train or fine tune as well, but the first step is just being able to keep the weights in memory. Inference isn't that computationally intensive relative to just how much RAM you need.
- brucethemoose2 3y agoYeah. I didn't mention them because the 64GB+ configs are very expensive, and I don't think you can even finetune on them.
- Me1000 3y agoThey're not cheap, but they're not _that_ expensive compared to buying four NVLink'd Nvidia cards with a combined similar amount of VRAM. Plus you get a whole computer with it and you don't have to worry about casings and power supplies, etc. But yeah, you will be more compute constrained if you go down that route.
- brucethemoose2 3y agoThe price is similar to 2x RTX 8000s, or even A6000s, but yeah your point stands. Power efficiency is something too. You run into immense pain the moment you venture outside of llama inference though.
- brucethemoose2 3y agoVRAM is everything. The more VRAM you have, less aggressively you have to quantize models for inference, which in turn has huge speed/quality implications. You can run higher batch sizes, or draft models, or more caching, which increases efficiency. For LLMs specifically, you can load bigger models into VRAM in the first place. You can load more of a multimodal pipeline in VRAM without having to constantly swap everything out. This is all 10x true for finetuning. Quality and speed is essentially determined by VRAM capacity, as long as you are not on truly ancient GPU like a P40 than't can't even do fp16. As for architecture... TBH, many operations are heavily bandwidth bound these days. Sometimes a 3090 and a 4090 are essentially the same speed. And waiting a little longer for a finetune is no big deal vs not being able to do it at all, or doing it at low quality. > They are graphical cards, equipped with consumer grade VRAM (instead of HBM) and IMHO it will take a very big shift before we see them being used as real AI accelerators I think its important for users to break away from the cloud and APIs, and try to run stuff themself, lest we get locked into an OpenAI monopoly. But setting that aside, running local is also extremely useful for prototyping and testing. You can see if something works without burning dollars every second you spend debugging on a big cloud instance. Even if that's affordable, just feeling like I am under the clock when debugging/optimizing is stressful to me.