3 ms·
Are you just talking about transformers? OpenAI utilized them, yes, but they are the ones who authored the original GPT paper. I am not sure I follow.
by maxdoop 4y ago
Are you just talking about transformers? OpenAI utilized them, yes, but they are the ones who authored the original GPT paper. I am not sure I follow.
- lobstersammich 4y agoI think they mean Transformers in the Vaswani et al 'Attention is all you need' paper, not Generative Pretrained Transformers, specifically? Paper link below: https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf https://proceedings.neurips.cc/paper/2017/file/3f5ee243547de... For some papers on attention mechanisms from before the 2017 'Attention is all you need' paper, check out that paper's references. Chris Manning's 2015 paper covers attention mechanisms. And so do a few other researchers from that mid-2010s time period: [21] Minh-Thang Luong, Hieu Pham, and Christopher D Manning. Effective approaches to attention based neural machine translation. arXiv preprint arXiv:1508.04025, 2015. [22] Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit. A decomposable attention model. In Empirical Methods in Natural Language Processing, 2016.
- gcr 4y agoMy understanding is that the original Transformer paper ("Attention is All You Need") and BERT papers came from Google. In particular, BERT's pre-training task is to predict how to fill-in-the-blank from a sentence that's missing a random word. OpenAI started from that architecture and made it output text generatively, by using beam search on top of a next-token prediction task.