3 ms·
!Long post warning! Tokenization is often scapegoated for many transformer limitations. I suppose it's because reading about the many limitations of the transf
by Vetch 2y ago
!Long post warning!
Tokenization is often scapegoated for many transformer limitations. I suppose it's because reading about the many limitations of the transformer architecture is harder than dumping everything on tokenization (which to be fair, is often indirectly involved with or exacerbating some deeper issue).
> Why can't LLM spell words? Tokenization.
LLMs can spell if you ask them to though. And there have been investigations into this capability (ref:2). Tokenization makes computations that involve spelling more difficult, but this is downstream of deeper computational limitations of the architecture.
> Why can't LLM do super simple string processing tasks like reversing a string?
Ditto.
> Why is LLM worse at non-English languages (e.g. Japanese)? Tokenization.
Tokenization is also implicitly performing compression. If your tokenizer's corpus is focused only on english, basic information theory explains why it'll be less efficient for other languages. The net effect is longer sequences where tokens are less information dense for non-english languages on average.
> Why is LLM bad at simple arithmetic? Tokenization.
Tokenization could treat digits separately and I believe, llama2 did this. But OpenAI built tiktoken which does not do this. llama3 uses tiktoken.
The transformer architecture also has limitations that make (default) arithmetic computations involving carries difficult to learn. You can read more about this in (ref:1).
> Why did my LLM abruptly halt when it sees the string "<|endoftext|>"? Tokenization.
Why should it not? Either way, it doesn't have to halt, as the sampler can just ignore this. But the distribution will still condition on this as a change of topic switch. The question should probably be, why did the LLM suddenly assign high probability to a stop token before finishing whatever it was writing?
> What is this weird warning I get about a "trailing whitespace"? Tokenization.
Modeling decisions for how to treat whitespace is upstream of tokenization. These choices affect how the LLM models word boundaries. Things can be fine most of the time until they aren't.
There's also the issue of softmax. The way softmax is typically applied forces the model to always assign importance to some tokens, even when no strong relationships exist between them. This in turn leads to the model disproportionately dumping its focus on often semantically unimportant tokens like whitespace or punctuation. Misallocating attention in this manner can lead to wasting representational capacity due to overemphasizing unimportant tokens, perhaps inducing spurious correlations on whitespace. This issue propagates through the model, possibly leading to unexpected negative downstream effects.
> Why the LLM break if I ask it about "SolidGoldMagikarp"? Tokenization.
One step down, it's really a result of high dimensional random vectors.
> Why should I prefer to use YAML over JSON with LLMs? Tokenization.
> Why did GPT-2 have more than necessary trouble coding in Python? Tokenization.
Tokenization does make counting more difficult but the net benefit to programming languages where whitespace can be semantically meaningful is a strong positive. Even when whitespace is not meaningful, long strings of them can often be encountered. Not being careful about devoting tokenization effort on whitespace will significantly degrade code modeling ability in LLMs.
> Why is LLM not actually end-to-end language modeling? Tokenization.
This is correct, but it is not necessarily the case that a character or byte based model will automatically be better. The issue is that LLMs as currently devised spend the same amount of computation per token. This creates the immediate problem of making meaningful sequences, which will now be substantially longer, substantially more expensive to compute, generate and store in memory. This is what the posted paper seeks to address over naive byte level modeling. Although it's unclear from the provided tables if what's claimed is actually what's occurring.
Character level modeling will also make learning long ranged dependencies harder. Subword tokenization also aids in memorization, which can be useful in learning from the tail of the distribution. The following idea is based on (ref:5).
Next-token prediction can be modeled as a hierarchical sampling process where problem instances (topics, natural language tasks), which are mixture distributions, are drawn from a metadistribution, and then data points (eg various strings) are sampled from specific subpopulations (ie clusters of task types) within those instances. Here, memorization is a key strategy since there's initial uncertainty about which features are relevant for predicting the next token. Particularly for rare examples, memorizing their details acts as a starting point for associating particular patterns with specific subpopulations, in turn allowing more accurate prediction of new points.
From that starting point, the model can eventually refine its associations as it encounters more data. This is key for example, when sampling from the tail of the distribution where data about subpopulations will be more limited. Making memorization and learning longer dependencies more challenging can lead to final models that face more difficulty during ICL inference, which depends, among other things, on the ability to infer which task from a mixture distribution.
> What is the real root of suffering? Tokenization.
A better candidate is over-generalization.
1: https://arxiv.org/abs/2310.16028 https://arxiv.org/abs/2310.16028
2: What do tokens know about their characters and how do they know it? (https://aclanthology.org/2022.naacl- https://aclanthology.org/2022.naacl-
main.179.pdf)
3: https://arxiv.org/abs/2406.10851 https://arxiv.org/abs/2406.10851
4: Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP (https://arxiv.org/abs/2112.10508 https://arxiv.org/abs/2112.10508)
5: https://arxiv.org/abs/2012.06421 https://arxiv.org/abs/2012.06421