4 ms·
Positional description matters for transformers arithmetic
- thomasahle 3y agoAlso positional encodings. A few months ago I did a bunch of experiments with length generalization in addition and multiplication (https://github.com/thomasahle/arithmetic-transformer https://github.com/thomasahle/arithmetic-transformer). It turned out that not using positional encodings at all (relying entirely on the causality mapping in the attention) did better than all other methods like Learned PoE, Sinusoidal, RoPE and using an LSTM below the transformer.
- inciampati 3y agoFascinating, thanks for sharing this!
- Kerbonut 3y agoCan that be part of a multimodal language model to better understand and perform arithmetic operations, enhancing the LLMs numerical reasoning?