4 ms·
> Later decoder-only Transformer was shown to achieve great performance in language modeling tasks, like in GPT and BERT. Actually, BERT is an encoder-only arc
by spi 4y ago
> Later decoder-only Transformer was shown to achieve great performance in language modeling tasks, like in GPT and BERT.
Actually, BERT is an encoder-only architecture, not decoder-only. Aside from trying to solve the same problem, GPT and BERT are quite different. This kind of confusion on now "classic" transformer models makes me kind of dubitative that the more recent and exotic ones are described very accurately...
(Clicking on the link with more details on BERT actually doesn't dispel much of the confusion; it stresses the fact that unlike GPT it's bidirectional, and indeed bidirectional is the "B" in BERT, but that's quite a disingenuous choice of terms itself - it's not "bidirectional" as in Bi-LSTM, that go left-to-right and right-to-left separately, it does the whole sequence at once; that was the real innovation of BERT).
Scrolling down to Transformer-XL starts talking about segments, from the context I _think_ it means that the input text is split into segments that are dealt with separately to cut down on the O(N^2) dependency of the transformer, but I would have assumed this kind of information to be written in a survey article.
IMHO, review articles are really great and useful, because they allow to cut through the BS that every paper has to add to get published, unify notations, and summarize the main points clearly. This article does a commendable job on the second point and, partly, on the first, but sadly lacks the third. Given the enormous task that it certainly was to compile this list, it would probably have profited from treating fewer models but putting things a bit more into perspective...