Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
SoerenL
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
SoerenL
8y ago
Yeah, I am referring to Ricci notation. TF and PyTorch don't use it. The new version already has relus. This version is targeted at standard formulas (has abs as a non-differentiable function), next version works for deep learning. Nev
2.
▲
by
SoerenL
8y ago
True, for large-scale problems Hessian-vector products are often the way to go (or completely ignoring second order information). However, computing first an expression for the Hessian symbolically and then taking the product with a vector
3.
▲
by
SoerenL
8y ago
Not in its current formulation. It uses a different representation of the tensors. However, a new version/algorithm that will be available in a few months can be used in TF and PyTorch.
4.
▲
by
SoerenL
8y ago
That's what is usually done, autodiff on the component wise expression. We don't do it here. Instead, we really compute on the matrix and tensor level and compute derivatives here directly. Let me give you a simple example to illu
5.
▲
by
SoerenL
8y ago
Do you need Hessians or Jacobians for computing the earth mover distance, then a definite yes. Otherwise, I would doubt it (though I do not know exactly.)
6.
▲
by
SoerenL
8y ago
As one of the authors: Yes, it is true, 3 orders of magnitude (so about a factor of 1000 on GPUs). But please be careful, this holds only for higher order derivatives (like Hessians, or Jacobians), as stated in the paper and as the title sa
7.
▲
by
SoerenL
9y ago
But how do you compute the derivative of x' A x in Mathematica (x being a vector and A being a matrix)? What you have pointed out is only scalar derivatives, if I am not mistaken here.
8.
▲
by
SoerenL
9y ago
XLA is good for the GPU only. On the CPU MC is about 20-50% faster than TF on scalar valued functions. For the GPU I don't know yet. But it is true that for augmented Lagrangian you only need scalar valued functions. This is really eff
9.
▲
by
SoerenL
9y ago
I did not compare to Tensorflow XLA but I compared it to Tensorflow. Of course, it depends on the problem. For instance, for evaluating the Hessian of x' A x MC is a factor of 100 faster than TF. But MC and TF have different objectives
10.
▲
by
SoerenL
9y ago
As one of the authors of this tool I understand that you would like to have some way of trusting the output. Internally we check it numerically, i.e., generate some random data for the given variables and check the derivative by comparing i
11.
▲
by
SoerenL
9y ago
It is not in the current online tool but we will add it again soon. It is still in there the way you describe it (passing transpose down to the leaves and simplification rules as well). Btw: How are you doing and where have you been? Would