6 ms·
Side question: Can people this deep in the field read that visualization with all the formulas and actually grok what's going on? I'm trying to understand just
by byteknight 2y ago
Side question:
Can people this deep in the field read that visualization with all the formulas and actually grok what's going on? I'm trying to understand just how far behind I am from the average math person (obviously very very very far, but quantifiable lol)
- Mc91 2y agoI'm not deep in the field at all, I did about four hours of Andrew Ng's deep learning course, and have played around a little bit with Pytorch and Python (although more to install LLMs and Stable Diffusion than to do Pytorch directly, although I did that a little too). I also did a little more reading and playing with it all, but not that much. Do I understand the Python? Somewhat. I know a relu is a rectified linear unit, which is a type of activation function. I have seen einsum before but forget what it is. For the classical diagram I know what the nodes, edges and weights are. I have some idea what the formulas do, but not totally. I'm unfamiliar with tensor diagrams. So I have very little knowledge of this field, and I have a decent grasp of some of what it means, a vague grasp on other parts, and tensor diagrams I have little to no familiarity with.
- thomasahle 2y agoThe tensor diagrams are not quite standard (yet). That's why I also include more "classical" neural network diagrams next to them. I've recently been working on a library for doing automatic manipulation and differentiation of tensor diagrams (https://github.com/thomasahle/tensorgrad https://github.com/thomasahle/tensorgrad), and to me they are clearly a cleaner notation. For a beautiful introduction to tensor networks, see also Jordan Taylor's blog post (https://www.lesswrong.com/posts/BQKKQiBmc63fwjDrj/graphical-tensor-notation-for-interpretability https://www.lesswrong.com/posts/BQKKQiBmc63fwjDrj/graphical-...)
- programjames 2y agoThese remind me of interaction combinators [1], which are being used in the Bend programming language [1]. I think it'd be good for the standard to also be a valid interaction net. [1]: https://core.ac.uk/download/pdf/81113716.pdf https://core.ac.uk/download/pdf/81113716.pdf [2]: https://news.ycombinator.com/item?id=40390287 https://news.ycombinator.com/item?id=40390287
- thomasahle 2y agoThis stuff is super cool! It basically generalizes tensor diagrams to general computational graphs. However, when thinking about ML architectures, I actually like that classical tensor diagrams make it harder to express non-associative architectures. E.g. RNNs are much harder to write than Transformers.
- cgadski 2y agoAfter learning about tensor diagrams a few months ago, they're my default notation for tensors. I liked your chart and also Jordan Taylor's diagram for multi-head attention. Some notes for other readers seeing this for the first time: My favorite property of that these diagrams is that they make it easy to re-interpret a multilinear expression as a multilinear function of any of its variables. For example, in standard matrix notation you'd write x^T A x to get a quadratic form with respect to the variable x. I think most people read this either left to right or right to left: take a matrix-vector product, and then take an inner product between vectors. Tensor notation is more like prolog: the diagram x - A - x involves these two indices/variables (the lines) "bound" by three tensors/relations (A and two copies of x.) That framing makes it easier to think about the expression as a function of A: it's just a "Frobenius inner product" between -A- and the tensor product -x x-. The same thing happens with the inner product between a signal and a convolution of two other signals. In standard notation it might take a little thought to remember how to differentiate <x, y * z> with respect to y (<x, y * z> = <y, x * z'> where x' is a time-reversal), but thinking with a tensor diagram reminds you to focus on the relation x = y + z (a 3-dimensional tensor) constraining the indices x, y and z of your three signals. All of this becomes increasingly critical when you have more indices involved. For example, how can you write the flattened matrix vec(AX + XB) as a matrix-vector product of vec(X) so we can solve the equation AX + XB = C? (Example stolen from your book.) I still have to get a hold of all the rules for dealing with non-linearities ("bubbles") though. I'll have to take a look at your tensor cookbook :) I'm also sad that I can't write tensor diagrams easily in my digital notes. Tensor diagrams are algebraically the same thing as factor graphs in probability theory. (Tensors correspond to factors and indices correspond to variables.) The only difference is that factors in probability theory need to be non-negative. You can define a contraction over indices for tensors taking values in any semiring though. The max-plus semiring gives you maximum log-likelihood problems, and so on.
- thomasahle 2y agoI'm really glad you've found my "book" useful! Makes me want to continue writing it :)
- cshimmin 2y agoI'm familiar with almost all of these architectures, but not the tensor diagram notation. I can't figure out what "B" is? I thought maybe it's a bias vector, but then why does it only appear on the input data, and not on subsequent fc layers?
- cgadski 2y agoB is the number of data vectors going on. You can erase the line labeled by B without much loss. (You just get the diagram for the feed-forward of a single vector.)
- Krei-se 2y agoYou don't need to be more good in math than in high school. AI is a chain of functions and you derive over those to get to the loss-function (gradient) to tell you which parameters to change to get a better result (simplified!). Now this structure of functions is different in each implementations, but the type of function is quite similar - even though a large model will combine billions of those nodes and weights. Those visualizations tell you f.e. that some models connect neurons back to ones earlier in the chain to better remember a state. But the activation function is usually a weight and threshold. KAN changes the functions on the edges to more sophisticated ones than just "multiply by 0.x" and uses known physical formulas that you can actually explain to a human instead of the result coming from 100x different weights which tell you nothing. The language models we use currently may map how your brain works, but how strong the neurons are connected and to which others does not tell you anything. Instead a computer can chain different functions like you would chain a normal work task and explain each step to you / combine those learned routines on different tasks. I am by no means an expert in this field, but i do a lot of category theory, especially for the reason that i wanted a more explainable neuron network. So take my pov with a grain of salt, but please don't be discouraged to learn this. If you can program a little and remember some calculus you can definitely grasp these concepts after learning the vocabulary!
- godelski 2y ago> You don't need to be more good in math than in high school. I'm very tired of this... it needs to stop as it literally hinders ML progress 1) I know one (ONE) person who took multivariate calculus in high school. They did so by going to the local community college. I know zero people who took linear algebra. I just checked the listing of my old high school. Over a decade later neither multivariate calculus nor linear algebra is offered. 2) There's something I like to tell my students You don't need math to train a good model, but you do need to know math to know why your model is wrong. I'm sure many here recognize the reference[0], but being able to make a model that performs successfully on a test set[1] is not always meaningful. For example, about a year ago I was working a very big tech firm and increased their model's capacity on customer data by over 200% with a model that performed worse on their "test set". No additional data was used, nor did I make any changes to the architecture. Figure that out without math. (note, I was able to predict poor generalization performance PRIOR to my changes and accurately predict my model's significantly higher generalization performance) 3) Math isn't just writing calculations down. That's part of it -- a big part -- but the concepts are critical. And to truly understand those concepts, you at some point need to do these calculations. Because at the end of the day, math is a language[2]. 4) Just because the simplified view is not mathematically intensive does not mean math isn't important nor does it mean there isn't extremely complex mathematics under the hood. You're only explaining the mathematics in a simple way that is only about the updating process. There's a lot more to ML. And this should obviously be true since we consider them "black boxes"[3]. A lack of interpretability is not due to an immutable law, but due to our lack of understanding of a highly complex system. Yes, maybe each action in that system is simple, but if that meant the system as a whole was simple then I welcome you to develop a TOE for physics. Emergence is useful but also a pain in the ass[4]. [0] https://en.wikipedia.org/wiki/All_models_are_wrong https://en.wikipedia.org/wiki/All_models_are_wrong [1] For one, this is more accurately called a validation set. Test sets are held out. No more tuning. You're done. This is self-referential to my point. [2] If you want to fight me on this, at least demonstrate to me you have taken an abstract algebra course and understand ideals and rings. Even better if axioms and set theory. I accept other positions, but too many argue from the basis of physics without understanding the difference between a physics and physics. Just because math is the language of physics does not mean math (or even physics) is inherently an objective principle (physics is a model). [3] I hate this term. They are not black, but they are opaque. Which is to say that there is _some_ transparency. [4] I am using the term "emergence" in the way a physicist would, not what you've seen in an ML paper. Why? Well read point 4 again starting at footnote [3].
- danielmarkbruce 2y agoYes. But it's not difficult math in 99% of cases, it's just notation. It may as well be written in Japanese.
- canjobear 2y agoNot hard to understand. The visualization is more or less the computation graph that PyTorch builds up. And the einsum code is even clearer. There’s definitely a practice effect though. I know people who aren’t used to it will have their eyes glaze over when they read einsum notation.