4 ms·
I'll throw my hat in the ring. A transformer is a type of neural network that, like many networks before, is composed of two parts: the "encoder" that receives
by probably_wrong 3y ago
I'll throw my hat in the ring.
A transformer is a type of neural network that, like many networks before, is composed of two parts: the "encoder" that receives a text and builds an internal representation of what the text "means"[1], and the "decoder" that uses the internal representation built by the encoder to generate an output text. Let's say you want to translate the sentence "The train is arriving" to Spanish.
Both the encoder and decoder are built like Lego, with identical layers stacked on top of each other. The lowest lever of the encoder looks at the input text and identifies the role of individual words and how they interact with each other. This is passed to the layer above, which does the same but at a higher level. In our example it would be as if the first layer identified that "train" and "arrive" are important, then the second one identifies that "the train" and "is arriving" are core concepts, the third one links both concepts together, and so on.
All of these internal representations are then passed to the decoder (all of them, not just the last ones) which uses them to generate a single word, in this case "El". This word is then fed back to the decoder, that now needs to generate an appropriate continuation for "El", which in this case would be "tren". You repeat this procedure over and over until the transformer says "I'm done", hopefully having generated "El tren está llegando" in the process.
The attention mechanism already existed before transformers, typically coupled with an RNN. The key concept of the transformer was building an architecture that removed the RNN completely. The negative side is that it is a computationally inefficient architecture as there are plenty of n^2 operations on the length of the input [2]. Luckily for us, a bunch of companies started releasing for free giant models trained on lots of data, researchers learned how to "fine tune" them to specific tasks using way less data than what it would have taken to train from scratch, and transformers exploded in popularity.
[1] I use "mean" in quotes here because the transformer can only learn from word co-occurrences. It knows that "grass" and "green" go well together, but it doesn't have the data to properly say why. The paper "Climbing towards NLU" is a nice read if you care about the topic, but be aware that some people disagree with this point of view.
[2] The transformer is less efficient that an LSTM in the total number of operations but, simultaneously, it is easier to parallelize. If you are Google this is the kind of problem you can easily solve by throwing a data center or two at the problem.
- jimbokun 3y ago> The negative side is that it is a computationally inefficient architecture as there are plenty of n^2 operations on the length of the input Is this the reason for the limited token windows?
- probably_wrong 3y agoYes, kinda. The transformer doesn't have a mechanism for dynamically adjusting its input size, so you need to strike a balance between the window being big enough for practical purposes but also small enough that you can still train the network. Previous networks with RNNs could in theory receive inputs of arbitrary size, but in practice their performance decreased as the input got longer because they "forgot" the earlier input as they went on. The paper "Neural Machine Translation by Jointly Learning to Align and Translate" solved the forgetting problem by, you guessed it, adding attention to the model. Eventually people realized that attention was all you needed (ha!), removed the RNN, and here we are.