4 ms·
> See, in 2017, Google researchers published the article “Attention is all you need,” introducing the concept of the transformer and vastly improving the capabi
by deepsquirrelnet 4y ago
> See, in 2017, Google researchers published the article “Attention is all you need,” introducing the concept of the transformer and vastly improving the capabilities of machine learning models.
…
You may well ask, why did Google give this wonderful thing away freely? While big private research outfits have been criticized in the past for withholding their work, the trend over the last few years has been toward publishing.
I think that whether or not they had published this article, the transformer architecture would have been established within a year or two anyway. If you read the paper, they cite plenty of references to work already using self attention. As attention itself has no trainable parameters, it’s a short jump to add an mlp to create the “memory bank” it needs.
Likely they knew this, and thought it better to publish and receive the credit for their work than to let someone else claim it.