3 ms·
If you're new to learning about regression and want to use it, I would prioritize the statistical issues. There are lots of computational details you could read
by imurray 10y ago
If you're new to learning about regression and want to use it, I would prioritize the statistical issues. There are lots of computational details you could read about, like in this post. However in simple cases, fitting least-squares linear regression is one line of code:
w_fit = X \ yy; % Matlab/Octave, X is NxD, yy is Nx1
w_fit = np.linalg.lstsq(X, yy)[0] # NumPy equivalent
Knowing how to construct the "design matrix" X is the important bit. What are your features, how will they be pre-processed, and given what you've done, should you trust w_fit to be meaningful, or just useful for prediction?
If you've thrown in lots of features, hoping to get good predictions, you may want to do "regularization". The computation is no harder for simple versions. One way to do "L2 regularization" is to add some extra rows to X and yy before fitting the weights:
% Matlab/Octave... the NumPy is nearly as simple.
[N, D] = size(X);
lambda = 1; % regularization constant
X = [X; eye(D)*sqrt(lambda)];
y = [yy; zeros(D,1)]
Or you could use an actual stats package like R's libraries that will contain a lot more care. Again, you need to understand the statistics, for example how you should select lambda. Then you'll want to know how to diagnose your model and whether you should trust it.
It's fun to work out the least squares maths and implement these things from scratch. But it has been done, and it's not what matters for most applications.
EDIT: a discussion with good links on linear models and beyond from the other day https://news.ycombinator.com/item?id=12237998 https://news.ycombinator.com/item?id=12237998
- geezerjay 10y ago> If you're new to learning about regression and want to use it, I would prioritize the statistical issues. There is absolutely no need to take the statistical route in least squares regression. In fact, it needlessly complicates things. Interpreting least squares regression as the minimization of the sum of squares of the residual is a straight-forward way to understand the technique.
- imurray 10y agoI'm not sure what you mean, but I think we're talking past each other. Regression is a statistical term. If it's going to be used for an application involving data and real decisions, not understanding the statistical issues is irresponsible. That's not at odds with viewing it as a procedure to minimize square errors. To clarify: There are probabilistic models (involving Gaussians and stuff) where some interpretations suggest least-squares estimation. That wasn't the distinction I was trying to make. Statisticians who hate probabilistic models will still say that it's important to think about what you put in X and y before you do X\y, and to be careful about what you can conclude from the result.
- apathy 10y agoStatisticians who hate probabilistic models might want to find another line of work. Like, say, philosophy. You use the phrase "involving Gaussians and stuff" to "clarify," and you ask others for rigor? Physician, heal thyself! And please don't speak for statisticians, they are a peculiar bunch. Ask the same question of 2 statisticians and you may well end up with 3 answers.
- geezerjay 10y ago> Regression is a statistical term. If it's going to be used for an application involving data and real decisions, not understanding the statistical issues is irresponsible. No, it isn't. It's curve fitting. The goal is to fit a curve to the data. In any engineering application, no one cares what's the statistical interpretation. They only want to fit a curve to the data, and do it in an optimal way. Hence, you get the sum of squares of the difference between the curve and the sample point (i.e., the residual) and minimize it by determining the parameters which minimize the regression. That's it.
- makeset 10y agoThat's a particularly incompetent and irresponsible view of engineering. The key to modeling anything correctly is to be aware of exactly what assumptions your choice of model corresponds to and how well they match the real-world processes underlying your data. So you've fitted some curve, but why that curve? What are the implications of assuming linearity between your input and output domains? What have you assumed about the distribution of noise over your predictions? How about over your input features? How would you expect the bias and variance of your predictions to degrade if any of these assumptions no longer held? As an engineer, if you don't understand the statistics behind your models enough to answer to such questions, their real-world applications aren't going to go far beyond wishful garbage-in-garbage-out number crunching.
- geezerjay 10y ago> So you've fitted some curve, but why that curve? Because it's the curve that's obtained by that particular minimization criteria, which is the minimization of the L² norm. If some other criteria was used, or other approximation function, then the result would also be valid. It appears that you are unaware that essentially all engineering in general, and whole field of computational mechanics in particular, is founded on what can be described as curve fitting. Whether it's plain old least squares approximation (in particular, moving least squares) or other techniques focused on the minimization of some other norm (Galerkin-type methods, for instance) the basis is all the same. > What are the implications of assuming linearity between your input and output domains? In short, because analytic functions and Taylor's theorem exist. Is it that hard? > As an engineer, if you don't understand the statistics behind your models enough to answer to such questions, their real-world applications aren't going to go far beyond wishful garbage-in-garbage-out number crunching. You know nothing about engineering, and somehow you're assuming that everything can only be valid if it's interpreted as a statistical problem. This isn't true, and it ignores complete knowledge field in physics and mathematics.
- meotai 10y agoTo add, here's a good standard text to help you learn to "diagnose your model" for bias & classical assumption violations. And how to fix it. https://www.amazon.com/Introductory-Econometrics-Modern-Approach-Economics/dp/1111531048 https://www.amazon.com/Introductory-Econometrics-Modern-Appr...
- jimmyerf 10y agoEconometrician here. I really hate that linear regression is being explained in "Machine Learning" language. "Features", "design matrix", "prediction s".
- cossatot 10y agoAnd many in the sciences are confused by 'endogenous' and 'exogenous', but just say 'independent' and 'dependent'... Every field has its jargon and econometrics has no primacy here; no reason to have a strong emotional reaction to words you understand because they're not what you learned in college.
- bigger_cheese 10y ago>Then you'll want to know how to diagnose your model and whether you should trust it. Yes this x1000. Normally I rely on F Values to get an indication of how reliable the thing is - the package I use spits out a tonne of other information - Cook's D etc but most of the output is just noise to me at university all I was taught was T-test (I'm an engineer) and the (now retired) statistician at my work only gave me a very cursory explanation and he basically emphasized F value.