3 ms·
The latter one. The network is trained in full precision (this is required for the gradient calculation), but the weights are nudged towards the quantized value
by cpldcpu 2y ago
The latter one. The network is trained in full precision (this is required for the gradient calculation), but the weights are nudged towards the quantized values.