Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jkam
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
jkam
8y ago
What was the gift from Google?
2.
▲
by
jkam
8y ago
What kind of work are you referring to when you say higher-order SGD may _now_ be feasible for deep learning? I only find results that try to approximate second order information.
3.
▲
by
jkam
8y ago
Are you referring to the Ricci notation when you are saying it uses a different representation of the tensors? Do you also plan to add non-differentiable functions like relu?
4.
▲
Computing Higher Order Derivatives of Matrix and Tensor Expressions [pdf]
(matrixcalculus.org)
95 points
by
jkam
8y ago
|
20 comments
5.
▲
by
jkam
9y ago
I think XLA is trying to reduce the overhead introduced by backprop, meaning when you optimize the computational graph you might end up with an efficient calculation of the gradient (closer to the calculation you get with MC). Regarding non
6.
▲
by
jkam
9y ago
I'm good. Looking at things from the data angle now. But unfortunately no public page. You can link to the old one, if you want to. Have you compared against TensorFlow XLA?
7.
▲
by
jkam
9y ago
I wonder why they don't do it. In the tool I had running, we handled this by removing any transpose of a symmetric matrix (after propagating it before the leaves). Together with the simplification rule x + x -> 2*x for any x, you ge