11 ms·
This was trained to be run at FP8 with no quality loss.
by ancientworldnow 2y ago
This was trained to be run at FP8 with no quality loss.
- hislaziness 2y agoThe model description on huggingface says - Model size - 12.2B params, Tensor type - BF16. Is the Tensor type different from the training param size?