13 ms·
Transformers as Support Vector Machines
- sdenton4 3y ago[strike]Punk's[/strike] SVM's not dead! (More seriously, it's good to find inroads to better formal understanding of what's happening in these systems.)
- adamnemecek 3y agoAll machine learning is about finding hyperplanes.
- nologic01 3y agoThe large dimensionality seems to be what creates the need for heuristic designs rather than a generic approach
- adamnemecek 3y agoHyperplanes are the heuristics.
- alexmolas 3y agoHyperplanes is all you need
- quickthrower2 3y agoHyperplanation is all you need
- jpfed 3y agoSomeday I'm going to write a paper that achieves SOTA results with a nigh-incomprehensible mishmash of diverse techniques and title it "All You Need Considered Harmful".
- tudorw 3y agoand multi-dimensional topological manifolds, maybe :)
- revskill 3y agoWhat is hyperplane ?
- adamnemecek 3y agoFor 2d, a line, for 3d a plane, for nd a hyperplane. https://en.wikipedia.org/wiki/Hyperplane https://en.wikipedia.org/wiki/Hyperplane
- noduerme 3y agoit's a word (a made up word)
- MattPalmer1086 3y agoAll words are made up!
- noduerme 3y agoYeah, but only a few are made up to seem like terms of art designed to obfuscate their actual meaning; and usually prepending "hyper-" to something is a signal that a more clear description of the thing doesn't yet exist. Downvote away, fellas.
- adamnemecek 3y agohttps://en.wikipedia.org/wiki/Hyperplane https://en.wikipedia.org/wiki/Hyperplane
- shakow 3y agoA subspace of dimension n-1 of a n-dimensional vector space. It is an extension of the well-known concept of a 2d-plane in a 3d-space to nd-spaces.
- eru 3y agoYou could also describe a hyperplane as the set of solutions of a system of linear equations.
- mafribe 3y agoThis is wrong! The term hyperplane already assumes that the hypothesis space that your learning algorithm searches has some kind of dimension and is some variant of an Euclidean / vector space (and its generalisations). This is not the case for many forms of ML, for example grammar induction (where the hypothesis space is Chomsky-style grammars) or inductive logic programming (hypothesis space are Prolog (or similar) programs), or, more generally, program synthesis (where programs form the hypothesis space).
- adamnemecek 3y agoIt can also just be some sort of partitioning. I would be really surprised if there was no partitioning of some space.
- mafribe 3y agoNote that "some sort of partitioning" isn't a hyperplane. A partition is a set-theoretic concept. A hyperplane is (a generalisation of) a geometric concept, so has much more structure.
- adamnemecek 3y agoAlright how about coalgebra.
- abhinai 3y agoFully connected neural networks are hierarchies of logistic regression nodes. Transformers are networks of SVM nodes. I guess we can expect networks of other kinds of classifiers in the future. Perhaps networks of Decision Tree nodes? Mix and match?
- thomasahle 3y ago> Fully connected neural networks are hierarchies of logistic regression nodes. Only if you use softmax ss your activation function.
- bjornsing 3y agoYou mean sigmoid activation function?
- thomasahle 3y agoIf we are talking about "hierarchies of logistic regression nodes" we have to define how to extend logistic regression to multiple outputs. The most common approach is Multinomial logistic regression: https://en.wikipedia.org/wiki/Logistic_regression#Extensions https://en.wikipedia.org/wiki/Logistic_regression#Extensions . Other times sigmoid might be the right answer.
- mjburgess 3y agoNNs are decision trees anyway -- take any classification alg and rewind from its decision points into a disjunction of conditions. Or, maybe more clearly: imagine taking any classification algorithm and drawing the graph of all of its predictions across it's domain. Then just construct a decision tree which "draws splits" along the original alg's decision edges. Likewise, all ML is equivalent to a KNN parameterised on an averaging operation. Everything here is eqv to everything else. ML is just computing an expectation over a training dataset, weighted by the model parameters. The "value" comes from the (copyright laundering/) data. The only question is: can you find useful weights by which to control the expectation you're taking? Various ML approaches weight the training data differently. The most successful of the latest round of AI manages to compute weights across everything ever written --- hence more useful than naive KNN which wouldnt terminate on 1PB of text.
- regularfry 3y agoPractically speaking, does this give us anything interesting from an implementation perspective? My uneducated reading of this is that a single SVM layer is equivalent to the multiple steps in a transformer layer. I'm guessing it can't reduce the number of computations purely from an information theory argument, but doesn't it imply a radically simpler and easier to implement architecture?
- choeger 3y agoI am waiting for someone publishing the theoretical limits of these "AI" systems. They're certainly impressive language models - don't get me wrong on that. But every algorithm and every model has its limits. To know the limits turns their application from hype into engineering. And of course, the hype-sellers will try to keep that from happening as long as possible.
- ben_w 3y agoHype sellers, despite being annoying and noisy, are not the reason why it's hard to figure out the theoretical limits. To put it the form of a rhetorical question: many of these models are public, so why "wait" when you could do the research yourself?
- fl7305 3y ago> I am waiting for someone publishing the theoretical limits of these "AI" systems. > To know the limits turns their application from hype into engineering. It would be helpful to know how the models actually work under the hood. But we made very good use of metals for thousands of years before we understood things like atoms, chemical bonds, lattices, etc. Some engineering disciplines can be made up largely of empirical knowledge. Engineering to me is "make the things we want out of the things we have", and not necessarily "design based on complete scientific theories".
- riwsky 3y agoI, as a Real Engineer, REFUSE to use ChatGPT until we have a working theory of quantum gravity. Enough of this bullshit where no one knows the fundamentals of what they’re working with.
- noduerme 3y agoFuck, imagine how many doctoral theses I could've written every time I tweaked a few lines of code to try some abstract way of recombining outputs I didn't fully understand. I missed the boat. All this jargon is absolutely for show, though. Purely intended to create the impression that there's some kind of moat to the "discovery". There are much clearer ways to express "we fucked around with putting the outputs of this black box back into the inputs", but I guess that doesn't impress the rubes.
- ResearchCode 3y agoYou're not wrong. Applied ML articles are not worth reading.
- constantly 3y agoI wouldn’t go this far, applied ML articles are my favorite articles. If you’re in the arena, it’s good to see things that other people have done from a practical perspective so you can ape it in your own work or not give it further consideration.
- kristopolous 3y agoI really really wish this culture of expressing simple things in ornate ways would die. All it does is make knowledge less accessible
- awesomeMilou 3y agoPeople have to get their PhD's somehow... ;)
- tinco 3y agoIf the excuse is true and the "ornate" language really is a dense representation of information then it should be fairly trivial to have an LLM agent unsummarize it. There could be a webservice that offers a parallel track of layman's translations of any paper.
- sgt101 3y agoUniversal function approximator == universal function approximator
- eru 3y agoTuring machines can also be used as universal function approximators. But I'm not sure it makes sense to put them in the same category as the other two.
- bjornsing 3y agoI would love to put them in the same category as the other two. In fact I’ve spent quite a lot if time thinking about it / experimenting. Wouldn’t it be great if we could somehow train on data and get a small Turing machine instead of a huge neural network?
- moffkalast 3y agoI would expect it to result in a large and slow Turing machine instead of a small neural network.
- ethbr1 3y agoSo far, betting on RISC over CISC in terms of ultimate hardware performance has been a good bet.
- eru 3y agoBoth RISC and CISC are usually used in the context of describing Turing complete instruction sets. I'm not sure, it's relevant here? If you want to make a comparison in this flavour: Turing machines are a bit like CPUs in that they can execute arbitrary things in sequence. All the flavours of machine learning are more like GPUs: they do well with oodles of big, parallelisable matrix multiplications interspersed with some simple non-linear transformations.
- 3y ago
- westurner 3y agoSVMs are randomly initialized (with arbitrary priors) and then are deterministic. From "What Is the Random Seed on SVM Sklearn, and Why Does It Produce Different Results?" https://saturncloud.io/blog/what-is-the-random-seed-on-svm-sklearn-and-why-does-it-produce-different-results/ https://saturncloud.io/blog/what-is-the-random-seed-on-svm-s... : > When you train an SVM model in sklearn, the algorithm uses a random initialization of the model parameters. This is necessary to avoid getting stuck in a local minimum during the optimization process. > The random initialization is controlled by a parameter called the random seed. The random seed is a number that is used to initialize the random number generator. This ensures that the random initialization of the model parameters is consistent across different runs of the code From "Random Initialization For Neural Networks : A Thing Of The Past" (2018) https://towardsdatascience.com/random-initialization-for-neural-networks-a-thing-of-the-past-bfcdd806bf9e https://towardsdatascience.com/random-initialization-for-neu... : > Lets look at three ways to initialize the weights between the layers before we start the forward, backward propagation to find the optimum weights. > 1: zero initialization > 2: random initialization > 3: he-et-al initialization Deep learning: https://en.wikipedia.org/wiki/Deep_learning https://en.wikipedia.org/wiki/Deep_learning SVM: https://en.wikipedia.org/wiki/Support_vector_machine https://en.wikipedia.org/wiki/Support_vector_machine Is it guaranteed that SVMs converge upon a solution regardless of random seed?
- Dr_Birdbrain 3y agoAn SVM is a quadratic program, which is convex. This means that they should always converge and they should always converge to the same global optimum, regardless of initialization, as long as they are feasible, I.e. as long as the two classes can be separated by an SVM.
- wizzard0 3y agomy tldr: this explains 1) why huge models are important (so the gradient is high-dimensional enough to be monotonic) 2) why attention (aka connections, aka indirections) is trainable at all; and says nothing about why they might generalize the dataset
- deleted 3y ago[deleted]
- SomeoneFromCA 3y agoTransformers as voltage amplifiers.
- u320 3y agoTransformers as toy vehicles that can turn into robots.
- quickthrower2 3y agoAnd robots in disguise
- deleted 3y ago[deleted]
- porridgeraisin 3y ago[flagged]
- noduerme 3y agoYou missed the best part where they think they're coming for our jobs.
- pedrosorio 3y agoI regret to inform you, I don’t think it’s the same set of people. Writing a cute “Transformers are SVMs” paper and “building chatGPT” are not the same skillset.
- deleted 3y ago[deleted]
- r-zip 3y agoDid you read the abstract?
- hexo 3y agoIs this an April joke?
- quickthrower2 3y agoI would love an Andrej video on this
- sametoymak 3y agoI am one of the authors. The most critical aspect is that transformer is a "different kind of SVM". It solves an SVM that separates 'good' tokens within each input sequence from 'bad' tokens. This SVM serves as a good-token-selector and is inherently different from the traditional SVM which assigns a 0-1 label to inputs. This also explains how attention induces sparsity through softmax: 'Bad' tokens that fall on the wrong side of the SVM decision boundary are suppressed by the softmax function, while 'good' tokens are those that end up with non-zero softmax probabilities. It is also worth mentioning this SVM arises from the exponential nature of the softmax. The title of the paper does not make this clear but hopefully abstract does :).
- ogogmad 3y agoWhen you say SVM, do you mean any classifier that finds a separating hyperplane, like a no-hidden-layer "perceptron" or Naive Bayes, instead of one which finds the maximum margin hyperplane? Or is finding the maximum margin important here? Thanks. Very interesting. I think our own brains and nervous system use a step-function as their "activation function", so this could - optimistically - be a throwback to the roots of Rosenblatt's idea.
- sametoymak 3y agoThis SVM summarizes the training dynamics of the attention layer, so there is no hidden-layer. It operates on the token embeddings of that layer. Essentially, weights of the attention layer converge (in direction) to the maximum margin separator between the good vs bad tokens. Note that there is no label involved, instead you are separating the tokens based on their contribution to the training loss. We can formally assign a "score" of each token for 1-layer model but this is tricky to do for multilayer with MLP heads. Finally, I agree that this is more step-function like. There are caveats we discuss in the paper (i.e. how TF assigns continuous softmax probabilities over the selected tokens). To me, summary is: Through softmax-attention, transformer is running a "feature/token selection procedure". Thanks to softmax, we can obtain a clean SVM interpretation of max-margin token separation.
- visarga 3y ago
- gugagore 3y agoSVMs typically have weights per data point. I.e. nonparametric/hyper parametric. Modern machine learning doesn't really work like that anymore, right?
- exegete 3y agoYes SVM’s don’t store weights like parametric models but they also don’t store weights “per data point”. Only the points closest to the decision boundary are stored (i.e., the “support vectors”).
- mjhay 3y agoThe weight per datapoint thing is actually kind of orthogonal to the concept of an SVM, but is conflated by most introductions to SVMs. SVMs are linear models using hinge loss. In the "primal" optimization perspective (rather than the dual problem SVMs are usually formulated as), one optimizes the feature weights like normal. This is not sparse in general, but it's not like dual SVM weights are particularly sparse in practice.
- gugagore 3y agoTotally. Thank you for expanding on "typically". If I can expand on your "kind of", it would be that because of the kernel trick, it actually does matter that the data itself can determine the "linear" (in an infinite dimensional space, that would require infinitely many parameters under the primal formulation) model.
- mjhay 3y agoKernelization can be done in primal or dual. Due to the representation theorem, it only ever needs as many parameters as data points. In the primal with a kernel K, you're just doing a feature expansion where each data point x corresponds to a feature whose value at each data point y is just K(x, y).
- joaogui1 3y agoThe attention matrix is computed based on all tokens in the context, so it kind of functions non-parametrically (but over the batch instead of over the whole training dataset)