3 ms·
It has seemed to me that the GPT would be considerably better at numbers if it just considered each digit as a token. Has anyone actually done an experiment to
by thethirdone 3y ago
It has seemed to me that the GPT would be considerably better at numbers if it just considered each digit as a token. Has anyone actually done an experiment to test this?
I wouldn't disbelieve that the grouped version is actually better with data, but it fights my intuition pretty hard. Grouping based on frequency obfuscates the regular nature of numbers.
- famouswaffles 3y ago>It has seemed to me that the GPT would be considerably better at numbers if it just considered each digit as a token. Has anyone actually done an experiment to test this? Yeah Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs - https://arxiv.org/abs/2402.14903 https://arxiv.org/abs/2402.14903 xVal: A Continuous Number Encoding for Large Language Models - https://arxiv.org/abs/2310.02989 https://arxiv.org/abs/2310.02989 I believe there's another paper that demonstrates something like also for the likes of spelling, counting etc but i can't remember it.
- thethirdone 3y ago> Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs - https://arxiv.org/abs/2402.14903 https://arxiv.org/abs/2402.14903 Very interesting paper. It does make sense to me the R2L chunking would be better than L2R chunking. It doesn't actually study single digit tokenization. I am mostly interested in a direct comparison between an LLM wide tokenization vs single digit tokenization. It would be nice to see a direct comparison between similarly trained models. Otherwise it is very hard to get a definitive answer by comparing models with varying sizes, training time, and general strength. > xVal: A Continuous Number Encoding for Large Language Models - https://arxiv.org/abs/2310.02989 https://arxiv.org/abs/2310.02989 I have seen this paper before, but hadn't payed attention to the p10 vs p100 analysis. Its not clear that the findings would be relevant to an LLM like gtp4 though.
- throwawaymaths 3y ago> Grouping based on frequency obfuscates the regular nature of numbers. It's literally a lookup table to go from token to embedding. What would you expect as improved? At best, maybe cache coherency if groups of numbers are converted to embeddings sequentially... But embeddings are huge (e.g. 8kb for llama-2) so you're losing caches jumping around between non-contiguous numbers anyways.
- thethirdone 3y agoThe nature of numbers as `A10^(n+1) + B10^n` for digits `XXXABXXX` is a very important relationship for doing any arithmetic. As you tokenize strings of digits, you lose the position information within the token make more complicated relationships between tokens because the total number of token pairs increases. For example in order for a super simple model to learn 3 digit multiplication, it would need to see at least one example for each token in order to get ANY information about what number it represents. Alternatively, with single digits you only need an example where each position is present in each location. Obviously, we would hope to have plenty of data, but I would expect better generalization from models which need to rely on memorization less. Alternatively, I can see a few reason why grouped digits would be better, but they are more complicated reasons than the reason above so by Occam's Razor my intuition says single digits should be better.