4 ms·
My kingdom for renaming this paper to something like "Tensor Product Attention is a Memory-Efficient Approach for Long-Sequence Language Modeling"
by carbocation 2y ago
My kingdom for renaming this paper to something like "Tensor Product Attention is a Memory-Efficient Approach for Long-Sequence Language Modeling"
- Zacharias030 2y agoIf you don’t like the title, wait till you see this acronym: „… we introduce the Tensor ProducT ATTenTion Transformer (T6), a new model architecture…“
- imjonse 2y agoThere is a famous transformer model named T5 from Google, and also S4, S4 and S6 (Mamba) in the LLM space, so it is not unusual naming.
- black_puppydog 2y ago"... is all you need" isn't unusual either, and yet GGP isn't happy about it (and I understand why)
- svantana 2y agoYes, but T5 is at least a normal acronym: Text-To-Text Transfer Transformer (albeit a bit forced)
- TeMPOraL 2y agoThat it's not unusual tells us that too many researchers in the field are chasing citations and fame at the expense of doing quality work.
- ben_w 2y agoMm. That or all sharing a sense of humour/in-jokes: I'm sure I'm not the only one here who immediately thought of "GOTO is all you need" and "Attention considered harmful"
- deleted 2y ago[deleted]
- superjan 2y agoI propose T-POT (Tensor Product attentiOn Transformer)
- prometheon1 2y agoTPOT already exists in the ML field, it was a somewhat popular autoML package a few years ago if I remember correctly and still seems to be around: https://github.com/EpistasisLab/tpot2 https://github.com/EpistasisLab/tpot2