4 ms·
Any deeplearning expert here. Why Neural network can't compute a linear function Celsius to Fahrenheit 100% accurately. Is it data or is it something can be op
by codesternews 7y ago
Any deeplearning expert here. Why Neural network can't compute a linear function Celsius to Fahrenheit 100% accurately.
Is it data or is it something can be optimised.
```
celsius_q = np.array([-40, -10, 0, 8, 15, 22, 38], dtype=float)
fahrenheit_a = np.array([-40, 14, 32, 46, 59, 72, 100], dtype=float)
for i,c in enumerate(celsius_q):
print("{} degrees Celsius = {} degrees Fahrenheit".format(c, fahrenheit_a[i]))
l0 = tf.keras.layers.Dense(units=1, input_shape=[1])
model = tf.keras.Sequential([l0])
model.compile(loss='mean_squared_error',
optimizer=tf.keras.optimizers.Adam(0.1))
history = model.fit(celsius_q, fahrenheit_a, epochs=500, verbose=False)
print("Finished training the model")
print(model.predict([100.0]))
// it results 211.874 which is not 100% accurate (100×1.8+32=212)
```
What can be done to make this NN 100% accurate for simple linear equation
𝑓=1.8𝑐+32
https://colab.research.google.com/github/tensorflow/examples/blob/master/courses/udacity_intro_to_tensorflow_for_deep_learning/l02c01_celsius_to_fahrenheit.ipynb#scrollTo=Y2zTA-rDS5Xk https://colab.research.google.com/github/tensorflow/examples...
- xvedejas 7y agoIt's the optimization strategy. Optimizers for neural nets are designed to perform well for non-linear problems, which comes at the expense of not providing exact solutions for linear regressions. A general exact solution is not possible for nonlinear, over-specified problems (which is what neural nets are good at), so strategies different than those used in linear regression are necessary.
- deleted 7y ago[deleted]
- creato 7y agoYour data is suspiciously rounded off. Just doing linear regression on that data isn't going to give you a perfect fit either.
- Gibbon1 7y agoThere is that. And also what I learned in school which is doing linear regression using a function with more degrees of freedom than the data tends to generate garbage. It can match the data points exactly and then be wildly off between them.
- ska 7y agoA classic demonstration of similar effect - any set of N data points in a time series, e.g. (t,f(t)), can be fit by a N-1 order polynomial to pass through each point. So fit a high order poly to a set of points sampled (esp. with a little noise) from a low order poly. You'll get crazy oscillations, and outside the sampling area it will likely diverge fast. Now add a smoothness term and crank it up until you get more reasonable results - regularization. It's a simplification, but informative about some ML techniques.
- Gibbon1 7y agoOne of my labs some students curve fitted a sixth order poly onto five data points they collected. The process being measured was y = something something minus ln(x). It fit all five points exactly and smoothly and wildly diverged on either side. The professor was really amused.
- chestervonwinch 7y agoIt’s called https://en.m.wikipedia.org/wiki/Runge's_phenomenon https://en.m.wikipedia.org/wiki/Runge's_phenomenon and can be mitigated eg by non-uniform interpolation grids like chebychev nodes.
- lawrenceyan 7y agoIf you wanted 100% accuracy, technically infinite data points and training time would be required. You could likely get >99% accuracy within a few epochs of training though. With such a simple function you're trying to emulate, your model will very quickly converge. Neural networks are much better suited for distilling down and compressing very complex high dimensional data though anyways, and you really don't need to be using them for problems like this. It's completely overkill in addition to being very computationally inefficient. There's nothing wrong with just simply using linear regression. In many cases it's the right choice. In your toy problem case you coded above, you are effectively just doing linear regression, except you added in an Adam gradient descent optimizer instead of just doing least squares, which by the way would have been infinitely faster and immediately given you an answer.
- czr 7y ago* Fix the data. Right now the optimal coefficients on your data (using least-squares) are m=1.79794911, b=31.952525636156476, which yields 211.74743638 when predicting on 100. * Tune the hyperparameters. In particular, tune the learning rate. To quote the Deep Learning Book [0]: > The learning rate is perhaps the most important hyperparameter. If you have time to tune only one hyperparameter, tune the learning rate. It controls the effective capacity of the model in a more complicated way than other hyperparameters—the effective capacity of the model is highest when the learning rate is correct for the optimization problem, not when the learning rate is especially large or especially small. The following code will yield exactly 212 almost every run (using fixed data and a different choice of learning rate): ``` celsius_q = np.array([-40, -10, 0, 8, 15, 22, 38], dtype=float) fahrenheit_a = np.array([x * 1.8 + 32 for x in celsius_q], dtype=float) for i, c in enumerate(celsius_q): print("{} degrees Celsius = {} degrees Fahrenheit".format(c, fahrenheit_a[i])) l0 = tf.keras.layers.Dense(units=1, input_shape=[1]) model = tf.keras.Sequential([l0]) model.compile(loss='mean_squared_error', optimizer=tf.keras.optimizers.Adam(lr=1.0)) history = model.fit(celsius_q, fahrenheit_a, epochs=500, verbose=False) print("Finished training the model") print(model.predict([100.0])) ``` [0] https://www.deeplearningbook.org/contents/guidelines.html https://www.deeplearningbook.org/contents/guidelines.html
- tylerhou 7y agoCode blocks don't work on HN; you need to format all your code with spaces: celsius_q = np.array([-40, -10, 0, 8, 15, 22, 38], dtype=float) fahrenheit_a = np.array([x * 1.8 + 32 for x in celsius_q], dtype=float) for i, c in enumerate(celsius_q): print("{} degrees Celsius = {} degrees Fahrenheit".format(c, fahrenheit_a[i])) l0 = tf.keras.layers.Dense(units=1, input_shape=[1]) model = tf.keras.Sequential([l0]) model.compile(loss='mean_squared_error', optimizer=tf.keras.optimizers.Adam(lr=1.0)) history = model.fit(celsius_q, fahrenheit_a, epochs=500, verbose=False) print("Finished training the model") print(model.predict([100.0]))
- wodenokoto 7y agoSet your weights manually, and you'll see that the network can easily compute any linear function.
- emilfihlman 7y agoYour activation function needs to be linear to be able to achieve fully linear results. However, this kills the neural network as a general purpose computation device.