3 ms·
In theory the 2B model should be somewhere in between, whenever it gets fixed.
by Namidairo 4y ago
In theory the 2B model should be somewhere in between, whenever it gets fixed.
- moyix 4y agoAssuming that it's just that hardcoded check I should be able to add it today. But if that check is load-bearing (i.e. if it relies on those assumptions elsewhere in the code) it could be a bit more painful. Edit: sadly, it's not that easy. Removing that check lets the 2B model load, but it produces gibberish. I've opened an issue with FasterTransformer here, and will also try to debug it further myself, but unfortunately it's not obvious how they're using that assumption. https://github.com/NVIDIA/FasterTransformer/issues/268 https://github.com/NVIDIA/FasterTransformer/issues/268
- moyix 4y agoGot it working :) You can now use the 2B models in FauxPilot as well.