3 ms·
About the work you cite, I think that double descent is simply because the extra number of parameters (low bias) used produces a large variance when the input d
by manthideaal 7y ago
About the work you cite, I think that double descent is simply because the extra number of parameters (low bias) used produces a large variance when the input data is small, but as more data is introduced the extra parameters don't play any role, that is they are prunned. So the high bias is relative to the quantity of availabe information (training data). So the second descent starts when the extra parameters are prunned, in practice their coefficients tends to zero, the system learns that those coefficient don't play any role.