4 ms·
I did this with 24GB VRAM and 32GB DDR5, using LM Studio, and it was about as fast as I could read. (I read fast but I'd have to run it again to guess the toke
by EarthLaunch 3y ago
I did this with 24GB VRAM and 32GB DDR5, using LM Studio, and it was about as fast as I could read. (I read fast but I'd have to run it again to guess the token rate.)
I'm upgrading to 96GB RAM now to run the larger models, but I do wonder whether it'll be slow when using proportionally less VRAM.