3 ms·
We do quantization-aware training, so the model should minimize the loss w.r.t. the ternary weights, hence no degradation in performance.
by areddyyt 2y ago
We do quantization-aware training, so the model should minimize the loss w.r.t. the ternary weights, hence no degradation in performance.