3 ms·
Having every model re-trained in each language is a certain path towards having any non-English (or at most a couple of other languages from countries with big
by spi 3y ago
Having every model re-trained in each language is a certain path towards having any non-English (or at most a couple of other languages from countries with big pockets, like Chinese) language model be always massively behind - the resources required to train a model are huge, you can't expect e.g. the Polish community (plus anyone else) to replicate every good English model that comes out. GPT4 is less capable in Polish than in English, but probably much more than any Polish-specific model ever trained - and I suspect the gap is bigger than that with the best non-GPT4 English model.
Furthermore, I think you are exaggerating the memory issue of multilingual models significantly. Especially for languages using the same (Latin) script, the additional characters to care about are very few. Also a significant part of the vocabulary and language fall into a few buckets, so training a joint model makes all the sense in the world - much like an Italian native speaker could likely study a scientific text in Spanish and understand its content, even without speaking the language.
The memory impact comes mostly from having bigger embedding layers that have to account for vocabulary in many languages (the most problematic case being Chinese and Japanese, with their huge set of tokens). But even there, the largest vocabularies in use are maybe of size 100k (vs. about 30k for English-only), with a hidden dimension of 4k that makes for a total of 400M parameters. It's a lot, but a drop in the ocean of 100B+ parameters (or 1T+ for GPT4) we're seeing today.
P.S. Answering to GP, I think the Pile is English only, though - or at least, models on HuggingFace trained on the Pile, like the various Pythia models, are tagged as English only.