4 ms·
> A planned feature is to implement INT8 and INT4 support for the CodeGen models, which would let the models run with much less VRAM (~2x smaller for INT8 and ~
by Namidairo 4y ago
> A planned feature is to implement INT8 and INT4 support for the CodeGen models, which would let the models run with much less VRAM (~2x smaller for INT8 and ~4x for INT4) :)
That would certainly put the larger models within reach of most users. Part of my reluctance to deploy this myself was I'd have to either redline the VRAM on a 8GB card or choose the smallest model.
Would this require a retrain of the models?
- moyix 4y agoIt shouldn't require retraining, nope. I believe for INT4 there is a small adapter layer added that needs to be trained, but it is small and wouldn't require much data or computation to do so. Will know more once we actually start implementing :)