3 ms·
In this case I don't think what the article says is exactly right. Naively, tensors are just n-dimensional arrays, which TensorFlow supports. The paper linked i
by obastani 8y ago
In this case I don't think what the article says is exactly right. Naively, tensors are just n-dimensional arrays, which TensorFlow supports. The paper linked in the article appears more to be talking about how derivatives of tensors are represented in TensorFlow. The difference doesn't seem to matter unless you are taking higher-order derivatives. This makes sense, since TensorFlow is focused on first-order derivatives needed for gradient descent, but traditional machine learning algorithms also rely on second-order derivatives to make use of more powerful optimization algorithms based on Newton's method. I'm not sure exactly where the difference comes from, but it comes from a convenient notation for tensors used in physics, known as Einstein notation (Einstein invented this notation to make his life easier when deriving general relativity). In this notation, tensors are represented by a single scalar variable. For example, matrix multiplication y = A x is expressed as
y_i = A_ij x_j.
If I understand correctly, the paper points out that an algorithm for computing derivatives based on this notation is faster for taking higher-order derivatives compared to using TensorFlow.
Mathematically, tensors are more complicated objects. Basically, they are what you get when you take higher-order derivatives of a function. In particular, the first-order derivative of a function f: R^n -> R^m at a point x \in R^n is the best linear function A_x \in R^{m X n} that approximates the original function, i.e.,
f(x + dx) ~= f(x) + A_x dx.
A linear function is represented by a matrix, so a first-order derivative is a matrix. If I take the second-order derivative, I get a more complicated object B_x, which represents the quadratic term in the Taylor expansion:
f(x + dx) ~= f(x) + A_x dx + B_x(dx, dx)
where B_x(a, b) is a linear function (or more precisely, a "multilinear" function) of two vectors a, b (which are the same in the above formula). That is, whereas A_x is a (linear) function R^n -> R^m, B_x is a (multilinear) function R^n X R^n -> R^m. This mathematical object B_x is an example of a tensor. In R^n and R^m, tensors are pretty boring, but they become more interesting when dealing with functions on manifolds.
- max_likelihood 8y agoI have often seen Tensors introduced in the context of Einstein's General Relativity. I read this article on HN: https://news.ycombinator.com/item?id=19055994 https://news.ycombinator.com/item?id=19055994 a few weeks back and found it really helpful.
- sdenton4 8y ago+1. Looking quickly at the backing paper, it's all about higher-order derivatives. As I see it, the grindy-axe is about where and how one makes the hand-off from algebraic notation to actual computation: keeping the calculation in algebraic form allows efficient algebraic manipulations, which can then be translated into low-level computations. The question, then, is: a) whether the space of problems where you have good algebraic notation lines up well with the total scope of TF problems, and b) whether the extra complexity of supporting the full computer algebra system is 'worth it.' For the latter, keep in mind that algebraic derivatives can get cumbersome/expensive when you have an exponentially complex piecewise linear space (eg: https://arxiv.org/pdf/1711.02114.pdf); https://arxiv.org/pdf/1711.02114.pdf); the linked paper makes no mention of ReLUs... things might be fine with sigmoid activations, but they're the exception, these days...