3 ms·
The critique is about the importance of priors in BNN. In my humble understanding of Bayesian reasoning the argument to defend any prior is that with enought
by manthideaal 7y ago
The critique is about the importance of priors in BNN. In my humble understanding of Bayesian reasoning the argument to defend any prior is that with enought data the learning method converges to the real distribution, so if the result of any learning method depends heavily of any prior assumption then that assumption is crucial and in no way can it be taken randomly. On the other hand, it is well known that deep learning can learn any random model, so in the end I think all of this is about the bias-variance trade off. If your prior has an infinite number of adjustable parameters (zero bias) then the variance becomes infinite ( your result will depend and becomes equal to the training set).
So in practice one should choose the prior with the minimum number of parameters (bias) that shows a good learning performance on the available training set.
Anyway, trying to measure how a prior generalizes or not in BNN seems to be another way of thinking about bias-variance, if there is more than this, I would like to know.
- jacobbuckman 7y agoHave you seen the latest research on double descent? Here's a good intro, with references to some of the foundational work: https://openai.com/blog/deep-double-descent/ https://openai.com/blog/deep-double-descent/ It seems bias-variance doesn't apply to neural networks at all! So your intuitions are good, but there's definitely more to the story.
- manthideaal 7y agoAbout the work you cite, I think that double descent is simply because the extra number of parameters (low bias) used produces a large variance when the input data is small, but as more data is introduced the extra parameters don't play any role, that is they are prunned. So the high bias is relative to the quantity of availabe information (training data). So the second descent starts when the extra parameters are prunned, in practice their coefficients tends to zero, the system learns that those coefficient don't play any role.