3 ms·
I believe it comes down to the statistical frequency in the dataset used to build the tokenizer. If the dataset is 90% Burmese the resulting tokenizer will have
by snordgren 3y ago
I believe it comes down to the statistical frequency in the dataset used to build the tokenizer. If the dataset is 90% Burmese the resulting tokenizer will have single tokens representing common complex constructs in Burmese instead of English. Today's tokenizer might tokenize Burmese letter-by-letter or even byte-by-byte as it is not a very common language on the internet.
- jiggawatts 3y agoThe choice of tokenizer is arbitrary, and is configured by the people building the LLM. It's not dictated by the data. A simple model would be to simply tokenize the input character-by-character using Unicode as-is. Heck, you could use 256 (8-bit) tokens and just feed the model UTF-8 as a raw byte stream. The AI would "figure it out". However, it is much more efficient to use a tokenization tuned to the statistics of the input data set. Since most of the Internet is in English, it's more efficient to assign single tokens to entire English words, but not to other languages.
- astrange 3y ago> However, it is much more efficient to use a tokenization tuned to the statistics of the input data set. A program "tuned to the statistics of the input data set" is what an LLM is. So choosing to use a fixed tokenizer rather than letting the LLM learn one is a performance optimization, but not one we'll have to accept if models were designed better for performance themselves.
- bonzini 3y ago> it's more efficient to assign single tokens to entire English words Not entirely true, it's just that English has relatively limited inflection so words are not modified that much. In the weather example Spanish uses one more token than in English just because of the difference in sentence structure. The eight tokens are essentially "what weather will-be the week that come/s", i.e. the same sentence would use nine tokens in English. For a simpler example, Italian will use two tokens "legg/o" where English will also use two tokens but they will be entire words "I read". But "he read/s" may be three tokens where Italian uses two for "legg/e", because English in this case has some redundancy from its remaining vestiges of inflection. The fact that the tokenizers explored in the article aren't as efficient for non Latin alphabets is a different story.
- deleted 3y ago[deleted]