4 ms·
So I'm still wrapping my head around some of the math (I haven't had a math class in a handful of years)... I get the output of the model (y_model = m*xs[i]+b)
by JonathonCwik 10y ago
So I'm still wrapping my head around some of the math (I haven't had a math class in a handful of years)...
I get the output of the model (y_model = m*xs[i]+b), it's the y = mx + b where we know x (from the dataset) and have y be a variable.
The error is where I start to lose it, so I get the idea of the first part (ys[i]-y_model). It's basically the difference between the actual y value (from the dataset). I get that we want this number to be as small as possible as the closer to zero it is for the entire dataset that means we get closer to the line going through (or near) all the points and the closest fit will be when this total_error is nearest to zero.
What I don't get is the squaring of the difference. Is it just to make the difference a larger number so that it's a little more normalized? How do you get to the conclusion that it needs to be normalized? Same thing with the learning rate? I believe these to be correlated but I can't tell you how...
- jasode 10y ago>What I don't get is the squaring of the difference. The following answer is about standard deviation but the desired properties from squaring is also relevant to your question: http://stats.stackexchange.com/questions/118/why-square-the-difference-instead-of-taking-the-absolute-value-in-standard-devia http://stats.stackexchange.com/questions/118/why-square-the-...
- gkjohns 10y agoa) it ensures that they're positive so they don't just cancel each other out and b) like you mentioned, it penalizes huge errors more heavily. There are also historical reasons for using squared error. The square function is smooth and differentiable to you can analytically solve for the gradient. Before fast computers this was crucial for solving regression problems as a closed for makes everything easier.
- cowabungabruce 10y agoSquaring gets you guaranteed positive numbers. Remember, we are adding all the errors together to optimize the model: If we get sum_errors_A = 4 + -3 + -1 sum_errors_B = 1 + 1 + 1 B is obviously the better model, but it has a higher error than A when comparing. If we squared all the terms and then added, B would be the stronger model.
- JonathonCwik 10y agoAh, gotcha! Makes a lot more sense now.
- flor1s 10y agoI don't think there is much to gain from this tutorial, then again it doesn't pretend to offer you much either. For example in the code you are discussing, defining a model as variables and operations, instead of as a function, was confusing to me. Probably Tensorflow overloads the operations, but normally when you read "y = mx + b" you expect y to be computed directly and not to be stored as a model. "f = lambda m, x, b: m x + b" seems much more clear to me.