26 ms·
We will add quantized CodeGen for fast inference on CPUs up on cformers (https://github.com/NolanoOrg/cformers/ https://github.com/NolanoOrg/cformers/) by later
by ayushkaushal 4y ago
We will add quantized CodeGen for fast inference on CPUs up on cformers (https://github.com/NolanoOrg/cformers/ https://github.com/NolanoOrg/cformers/) by later today.
- underlines 4y ago4bit GPTQ maybe?
- syntaxing 4y agoWhoa is there a PR or wiki about this
- meghan_rain 4y ago> by later today Wow, that's the timeframe things are moving at right now, we better get used to it!