9 ms·
Machine Learning Crash Course: The Bias-Variance Dilemma
- Pogba666 9y agowow nice. Then I have things to do on my flight now.
- known 9y agoBrilliant post; Thank you;
- taeric 9y agoThis seems to ultimately come down to an idea that folks have a hard time shaking. It is entirely possible that you cannot recover the original signal using machine learning. This is, fundamentally, what separates this field from digital sampling. And this is not unique to machine learning, per se. https://fivethirtyeight.com/features/trump-noncitizen-voters/ https://fivethirtyeight.com/features/trump-noncitizen-voters... has a great widget that shows that as you get more data, you do not necessarily decrease inherent noise. In fact, it stays very constant. (Granted, this is in large because machine learning has most of its roots in statistics.) More explicitly, with ML, you are building probabilistic models. This is contrasted to most models folks are used to which are analytic models. That is, you run the calculations for an object moving across the field, and you get something within the measurement bounds that you expected. With a probabilistic model, you get something that is within the bounds of being in line with previous data you have collected. (None of this is to say this is a bad article. Just a bias to keep in mind as you are reading it. Hopefully, it helps you challenge it.)
- ssivark 9y agoI think the key point is that in the case of sampling, you always assume that your signal has finite bandwidth, so you can claim to avoid aliasing so long as you sample suitably often. Further, the time duration over which one makes measurements puts a cutoff on the frequency resolution. Both of these essentially serve the purpose of implicit model regularization (aka bias). For sufficiently well-behaved signals, the estimator of the strength of various frequency components (i.e. the Fourier transform) is pretty stable as one enlarges the window of measurement, which is post-hoc validation that the signal is well-behaved. This might not be true for very weird signals, and enlarging the window of measurement might significantly change model parameters--meaning that one might need a non-parametric model to get by (enlarging number of parameters with number of measurements) rather than restricting to a finite number of frequency models. Eg: Suppose the first 100 measurements show a nice sinusoid of some frequency, but the next 100 measurements show a pretty much flat signal. Then, you are forced to revise downward the parameter corresponding to the importance of the sinusoid component, and increase the importance of the zero frequency component. The thing is, one never truly knows how the signal is going to behave in the future. Any parametric model is a bias that things won't get too complicated.
- nightski 9y agoThe irreducible error is unrelated to the bias/variance trade-off. It's part of the overall error of course, but the bias/variance error is in addition to that. Unless I am misunderstanding your point.
- taeric 9y agoSorry, I did not mean to be that they were the same. And, to that point, I was greatly projecting based on my beliefs that have grown. The number one thing I keep having to re-stress and learn is that these are probabilistic models. So, if anything, I only meant they were related in that they are both good targets to internalize when working with ML. If there are better targets, I'm definitely interested in learning more.
- xapata 9y agoUltimately, everything is a probabilistic model. It'd be ridiculously impractical to try to incorporate all of one's knowledge into an analytic model. Instead, we generalize as appropriate and encapsulate the rest as a random variable, sometimes called "error". Plus, God plays dice and all that.
- cromd 9y agoIf it helps anyone, the FiveThirtyEight article describes a scenario where people take a survey about immigration status and voting. Most legal citizens will correctly identify themselves, but some will accidentally check the wrong box and say they are an illegal immigrant. If you have a billion citizens and 10 illegal immigrants truly taking the survey, and people check the wrong box 1 in 1000 times, your "percentage of illegal immigrants who vote" statistic will be about the same as for citizens (because almost all reported illegals will be citizens). Collecting more data won't help. It's a very good article, though in the context of deciding how many variables should be in a model of some complex phenomenon, this example is a little tougher to wrap your head around. It's not quite a predictive model, but there were some variables left out. A naive model I suppose is "this data is generated by infallible respondents", whereas a better model would incorporate that error rate. There isn't as much of a question about which pieces of information are relevant, though, like you might encounter when trying to predict future drug use from household income, race, age, number of books read as a child, number of pets, and so on.
- dopamean 9y agoIf you could find a link to the FiveThirtyEight article I'd really appreciate it. Thanks.
- foota 9y agoIt's linked in the grandparent's comment. https://fivethirtyeight.com/features/trump-noncitizen-voters/ https://fivethirtyeight.com/features/trump-noncitizen-voters...
- corey_moncure 9y agoDoesn't this make the assumption that "illegal immigrants" won't check the wrong box, intentionally or by accident?
- Double_Cast 9y agocitizens labeled illegals: 1,000,000 illegals labeled citizens: 0.01
- mljoe 9y agoFundamentally this is related to the induction fallacy. The doge meme image actually illustrates it pretty well. https://en.wikipedia.org/wiki/Problem_of_induction https://en.wikipedia.org/wiki/Problem_of_induction Another closely related thing is the No Free Lunch Theorem. https://en.wikipedia.org/wiki/No_free_lunch_in_search_and_optimization https://en.wikipedia.org/wiki/No_free_lunch_in_search_and_op... These are concepts I believe are very important to internalize if you work with machine learning. Fundamentally we are making predictions (ie. guessing) on the nature of entirely unknown information. So there is a certain inherent impossibility to the task in the general sense. It shouldn't always work.
- therajiv 9y agoWow, the discussion on the Fukushima civil engineering decision was pretty interesting. However, I find it surprising that the engineers simply overlooked the linearity of the law and used a nonlinear model. I wonder if there were any economic / other incentives at play, and the model shown was just used to justify the decision? Regardless, that post was a great read.
- pfd1986 9y agoMost likely, since building a facility to survive a 2.5x stronger shake would surely be a lot more expensive. I was also curious about how the data in the past few years did not follow the same trend as before. Does anyone know if that is what geologists call to be 'overdue' to an earthquake? Like California is supposed to be for a while?
- therajiv 9y agoWell, the data wasn't showing that the past few years were anomalous; rather, there were fewer high-magnitude earthquakes than expected. I don't think this has anything to do with being overdue for an earthquake. Most likely this is just because with events of low frequency (e.g. these higher-magnitude earthquakes were predicted to occur once every ~100 years by the linear model), large percent deviations from the expected value are more probable. Basically if you flip a coin 10 times you might imagine that 3 heads and 7 tails is pretty common, whereas 300 heads and 700 tails on 1000 tosses is comparitively extremely unlikely.
- pfd1986 9y agoMy point was that maybe for high mag quakes the power law is invalid... Or at least I dont think we have enough data at this end to be certain of what is going on. Here's another plot, this time from UK seismic frequency, where again the frequency for high magnitude earthquakes seem 'under' the expected curve. Yet, again, these are 2 plots... http://www.quakes.bgs.ac.uk/hazard/Hazard_UK.htm http://www.quakes.bgs.ac.uk/hazard/Hazard_UK.htm
- pfd1986 9y ago
- amelius 9y agoThe whole problem of overfitting or underfitting exists because you're not trying to understand the underlying model, but you're trying to "cheat" by inventing some formula that happens to work in most cases.
- pakl 9y agoMay I ask how you reached this insight? What field do you work in?
- platz 9y agoclassic statistics is much more interested in the explanatory power of models to describe phenomena. ML is mainly interested in prediction (correlation instead of causation), typically over some data that just fell in your lap.
- gaius 9y agoIn a sense Data Science is like the Cult of the MBA. MBAs believe a trained manager can manage anything because management skills are generic. A data scientist believes they can analyse anything because analysis is generic. Both fail in the real world because they discount domain knowledge.
- banned1 9y agoIs there a field that does not discount domain knowledge? Or is that just "judgment" and custom analysis? I am trying to understand how all fields map together. Thank you.
- sgt101 9y agoThe divisions are very confused. I think that sensible people all wish to use domain knowledge if possible. There are two tiers of this, firstly the use of domain knowledge in the manual or procedural construction of the insight system. Secondly the use of formalised knowledge in the creation of models that can then be fused with data. The first case is where data science has got a bad name; people swing into domains and companies full of cocksure ideas, produce insights that are risible or obvious and get ejected. Sometimes it takes years for sufficient knowledge to be acquired by analysts to deal with difficult domains. Lots of people use Bayesian inference to do the second. Tools like Stan and PyMC3 are really popular and effective.
- rdudekul 9y agoHere are parts 1, 2 & 3: Introduction, Regression/Classification, Cost Functions, and Gradient Descent: https://ml.berkeley.edu/blog/2016/11/06/tutorial-1/ https://ml.berkeley.edu/blog/2016/11/06/tutorial-1/ Perceptrons, Logistic Regression, and SVMs: https://ml.berkeley.edu/blog/2016/12/24/tutorial-2/ https://ml.berkeley.edu/blog/2016/12/24/tutorial-2/ Neural networks & Backpropagation: https://ml.berkeley.edu/blog/2017/02/04/tutorial-3/ https://ml.berkeley.edu/blog/2017/02/04/tutorial-3/
- plg 9y agolike many things in science and engineering, (and life in general) it comes down to this: what is signal, what is noise? most of the time there is no a priori way of determining this you come to the problem with your own assumptions (or you inherit them) and that guides you (or misguides you)
- eggie5 9y agoI've always liked this visualization of the Bias-Variance tradeoff: http://www.eggie5.com/110-bias-variance-tradeoff http://www.eggie5.com/110-bias-variance-tradeoff
- sandralopz565 9y agoIm making over $7k a month working part time. I kept hearing other people tell me how much money they can make online so I decided to look into it. Well, it was all true and has totally changed my life. This is what I do, ====http://bit.ly/2coUNgf http://bit.ly/2coUNgf
- culturedsystems 9y agoThat's OK as a visualisation of what bias and variance are, but it's a bad visualisation of the bias-variance tradeoff, because in that image there is no tradeoff - it shows a case where bias and variance are independent of one another. An illustration like this one genuinely confused me when I was first introduced to bias and variance: I couldn't understand why the lecturer was claiming there is a tradeoff while showing a picture of a case where there is no tradeoff. I eventually figured out what was going on, but I think I would have got it quicker if it had been explained more like the linked post, and less like that diagram.
- CuriouslyC 9y agoOne good way to solve the bias-variance problem is to use Gaussian processes (GPs). With GPs you build a probabilistic model of the covariance structure of your data. Locally complex, high variance models produce poor objective scores, so hyperparameter optimization favors "simpler" models. Even better, you can put priors on the parameters of your model and give it the full Bayesian treatment via MCMC. This avoids overfitting, and gives you information about how strongly your data specifies the model.
- gpawl 9y agoStatistics is the science of making decisions under uncertainty. It is far too frequently misunderstood as the science of making certainty from uncertainty.
- ehsquared 9y agoWelch Labs has a great 15-part series, where they gradually build up a decision tree model that counts the number of fingers in an image. Part 9 in the series explains the bias-variance spectrum really well: https://youtu.be/yLwZEuybaqE?list=PLiaHhY2iBX9ihLasvE8BKnS2Xg8AhY6iV https://youtu.be/yLwZEuybaqE?list=PLiaHhY2iBX9ihLasvE8BKnS2X...