4 ms·
Interesting, so they are using single precision (fp32), so 2.7B x 4 Bytes = ~10 GB. With CUDA overhead and room for context, you would need at least 12GB VRAM.
by armcat 3y ago
Interesting, so they are using single precision (fp32), so 2.7B x 4 Bytes = ~10 GB. With CUDA overhead and room for context, you would need at least 12GB VRAM. They could use half precision and half that VRAM requirement and save costs for everyone involved. Maybe there is a performance reason why they use full precision.
- collaborative 3y agoYes, the reason I asked is that when I see SLM I got all excited thinking "finally a small model that fits in cheap hardware for simpler tasks"
- eightysixfour 3y agoModels at this size are regularly quantized to 5/4bit to reduce the size. While there is some degradation, it isn’t as substantial as you would expect.
- deleted 3y ago[deleted]