5 ms·
For the benefit of laypeople like myself: - IID: "independent and identically distributed", https://en.wikipedia.org/wiki/Independent_and_identically_distribut
by EvilTerran 8y ago
For the benefit of laypeople like myself:
- IID: "independent and identically distributed", https://en.wikipedia.org/wiki/Independent_and_identically_distributed_random_variable https://en.wikipedia.org/wiki/Independent_and_identically_di...
- OLS: "ordinary least squares", https://en.wikipedia.org/wiki/Ordinary_least_squares https://en.wikipedia.org/wiki/Ordinary_least_squares (I think)
- notafraudster 8y agoYeah, to be clear, the discussion you're responding to is pretty out in the weeds. The great-grandparent to your content raised the imho well founded objection that many uses of machine learning "work" (produce good out-of-sample predictive accuracy) but we don't know "why". We have generated a distinct lack of theory. We know very little about assumptions and how they are violated. Separately, we know very little about why one technique works and another doesn't in a given context. We've also done very little thinking about how practitioners should deploy techniques except that predictive accuracy is good. The grandparent to your content was raising an objection that, actually, linear regression, a very old technique which in plain English means fitting a straight line to a scatterplot (but in any number of dimensions), has a great deal of theory around it. The simplest form of solving a linear regression is by "ordinary least squares" (minimizing the sum of squared deviations from the fit line): OLS. The grandparent was correct that especially in the mid-20th century to late 20th century, a lot of people did work on the conditions under which OLS works. What "works" means in a statistical sense is that it's efficient (has low uncertainty about the correct estimate), unbiased (on average gets the right answer), consistent (as you have more and more data gets closer to the right answer). Under a set of fairly impossible conditions about the real world data generated process, OLS is "BLUE" (the best linear unbiased estimator). Best here refers to efficiency, and unbiasedness I've already explained. OLS divides the data into structural elements (things that can be explained by the predictors you put into the estimator) and stochastic elements (the noise left over -- the deviations from the line). If we specify the correct model, the stochastic elements are the underlying stochasticity in the universe. If we specify an incorrect model, some of our omitted structural elements get put into the estimation of the noise. The grandparent noted that two assumptions made in linear regression are that the underlying stochastic disturbances in the data are i.i.d. (each is a random draw from the same distribution) and Gaussian (form a normal / bell curve). These are not assumptions, these are conditions under which OLS is BLUE. The latter is not a necessary condition at all, the distribution can take any form. The former is the most succinct way to express one of the conditions. My comment was to raise that actually when we use linear regression in the world, we rarely use classical OLS. In the real world underlying disturbances differ between observations. Imagine if I am running a regression on cross-country data, but while my US data is very precisely measured thanks to the widespread availability of polling firms, my Mexican data involves census enumerators going to rural villages. We might imagine that all of my US data is more precisely measured than all of my Mexican data, so we would expect the underlying stochasticity to differ between country. This is called clustering. Also, because we almost certainly do not have the correct model for the data (say log-dollars income predicts the result, not dollars income, but I put in dollars income), it can be the case that observations with higher values of our predicted Y also have more uncertainty. This is called heteroskedasticity. But the good news is we have answers to both, we just don't use OLS, we use more modern estimators. Yay! In general the world has moved away from rigorously teaching the conditions under which OLS and works and toward teaching more flexible estimators that work under less restrictive conditions. And in general in ML, people aren't using anything that looks anything like OLS, because ML has specific goals OLS is inappropriate for -- namely minimizing overfitting and out-of-sample error, where OLS is designed to maximize the precision of estimates of the slope parameters (how a given predictor affects the outcome). So all the work in OLS theory doesn't really translate to a machine learning setting, where many methods have no theory at all. Hope this is a plainer English version.
- EvilTerran 8y agoThat's a huge help, thank you!