3 ms·
What happens to this rift if it's operating on an equation with a derivative such as the gradient of a neural network? I'm a huge fan of deterministic computat
by scottlegrand2 7y ago
What happens to this rift if it's operating on an equation with a derivative such as the gradient of a neural network?
I'm a huge fan of deterministic computation, but we seem to be doing okay without it in deep learning with respect to floating point round off error. Or despite the fact that the computation could be deterministic, the algorithms are written in a way that they are not due to asynchronous accumulation and computation.
Would we just get a different minimum each time we run or would something more pernicious occur?