4 ms·
Isn't it better to adopt a Baynesian approach and directly model the complexity of the model rather than some arbitrary Ccomplexity? To wit, P(model|data)=P(da
by Robin_Message 16y ago
Isn't it better to adopt a Baynesian approach and directly model the complexity of the model rather than some arbitrary Ccomplexity?
To wit, P(model|data)=P(data|model)P(model) / P(data), which is to say that you calculate the probability a model is true given the data you see based on the probability you see the data you see given a certain model, multiplied by the probability of the model you are using.
I guess the problem in all this is we want* the model to be very complex, but we can't tell that complexity from overfitting (and it may be that no model exists so the only thing you can do is overfit.)
- alextp 16y agoYes, but the formulation I showed you is equivalent to a bayesian prior. For example, if you want to learn a weight vector w that gives high likelihood to the data and has a gaussian prior with 0 mean and Cidentity covariance, the MAP answer is "minimize -log(likelihood) + C||w||", where ||w|| is the square norm of w. Equivalently, if the prior is a laplacian you just change the norm from the l2 to the l1 norm. Being bayesian gives you an extra capability that is model averaging, and this does usually improve the behavior at a high computational cost. I really like bayesian models, and right now I'm experimenting with one that should do unsupervised sentiment analysis without a priori knowledge of word polarity or things like that (yes, I'm a phd student in machine learning).
- Robin_Message 16y agoI do remember reading that the Bayesian approach leads to previous empirical formulas falling out. Is that the case here, or was that formula derived using Bayes? I'm a PhD student in something else, and I'm trying to do some machine learning. So, what should I read to make what you said make sense :)?
- alextp 16y agoYes, it is the case here that a bayesian approach can lead to a previous empirical formula falling out. What I was saying as well is that this regularization + test set approach is also valid (and sometimes slightly more or less general than the bayesian approach, since, for example, SVMs fall naturally out of thinking about regularization but they have no analogue in bayesian classifiers). It also goes the other way, and some formulas are first proposed in a more bayesian-ish context and then extended to some simpler-looking empirical formulas (for example, the jumps from hidden Markov models to max-ent Markov models to conditional random fields to max-margin Markov networks to structured SVMs). There are more approaches to machine learning, and in John Langford's blog there is a nice table showing the merits and flaws of many of them: http://hunch.net/?p=224 http://hunch.net/?p=224 . But you must keep in mind that you can find many equivalencies between these approaches (boosting for example can be seen as a loss minimization with regularization, and max-ent can be seen as a special case of a bayesian model, etc).