3 ms·
Most LLMs determine their token inventories by using byte-pair encoding, which algorithmically induces sub-word tokens from a body of text. So even in English y
by lgessler 2y ago
Most LLMs determine their token inventories by using byte-pair encoding, which algorithmically induces sub-word tokens from a body of text. So even in English you might see a word like "proselytization" tokenized apart into "_pro", "selyt", "iz", "ation", and non-English languages will probably (depending on their proportional representation in the training corpus) also receive token allocations in the BPE vocabulary.
Here's actual output from the GPT-4o tokenizer for English and Hindi:
>>> [enc.decode([x]) for x in enc.encode("proselytization")]
['pros', 'ely', 't', 'ization']
>>> [enc.decode([x]) for x in enc.encode("पर्यावरणवाद")]
['पर', '्य', 'ावरण', 'वाद']