4 ms·
Without context, it should be number of params * 1.76 (the “effective bits per weight”) / 8 So for this one, 27B * 1.76 / 8 = 5.94 GB For speed, far as I can
by kennywinker 16d ago
Without context, it should be number of params * 1.76 (the “effective bits per weight”) / 8
So for this one, 27B * 1.76 / 8 = 5.94 GB
For speed, far as I can tell it depends if your gpu is memory bandwidth bound or (mostly older gpus) processing bound. If it’s memory bandwidth bound, and your gpu gets 300GB/s, that’s:
300GB/s / 5.98 GB = 50.5t/s.
Realistically it’s probably a bit slower, but that is your theoretical maximum.
- Dwedit 16d agoIt needs more VRAM than just the model weights. With 6GB of VRAM, I got 44/65 layers loaded into VRAM. Has anyone tested 8GB yet?
- kennywinker 16d ago“Without context” was me gesturing at that. Interesting you can’t fit the whole model tho - why is beyond my current understanding :)
- Dwedit 16d agoEven with a small context size (4096 at 64KB per token), that's like 256MB for the context. It's more than just the context that's eating up VRAM.