4 ms·
In my experience with a 4070ti (12 gb vram), 7800x3d and 64gb ddr5 it will run… but seemingly at about 0.1-0.5 tokens per second with partial GPU offload. It’s
by transcriptase 2y ago
In my experience with a 4070ti (12 gb vram), 7800x3d and 64gb ddr5 it will run… but seemingly at about 0.1-0.5 tokens per second with partial GPU offload.
It’s painful, but the results are night and day compared to the models that will fit entirely in vram. Of course perhaps I’m doing something wrong so if others have advice it would be great to hear!
- hnuser123456 2y agoIs it any faster with no GPU offload, given how little would fit in VRAM? Can you try giving a few prompts and getting a precise average token speed? Curious if it would be worth upgrading my RAM to at least run it on CPU. I have a 9800x3d/48gb ddr5/3090 and this model still being ~40GB when quantized is a challenge. Sure would be nice to have a 30B version that can quantize down to ~16GB. Or I guess I could upgrade my PSU and get another 3090 and use NVlink...