3 ms·
It's actually the other way round - the Transformer architecture was introduced for text (machine translation) in "Attention Is All You Need" (2017). Vision Tra
by laruss5 29d ago
It's actually the other way round - the Transformer architecture was introduced for text (machine translation) in "Attention Is All You Need" (2017). Vision Transformers, which apply it to images, came three years later in 2020: https://arxiv.org/abs/2010.11929 https://arxiv.org/abs/2010.11929
- verdverm 28d agoright, it was not text generation per-se (completion/contemporary understanding) that came first, vision was before that, translation before that