4 ms·
> Mistral NeMo uses a new tokenizer, Tekken, based on Tiktoken, that was trained on over more than 100 languages, and compresses natural language text and sourc
by PoignardAzur 2y ago
> Mistral NeMo uses a new tokenizer, Tekken, based on Tiktoken, that was trained on over more than 100 languages, and compresses natural language text and source code more efficiently than the SentencePiece tokenizer used in previous Mistral models.
From Mistral's page about Tekken:
> Our newest tokenizer, tekken, uses the Byte-Pair Encoding (BPE) with Tiktoken.
Does that mean that Mistral found that BPE is more efficient than unigram models?
Because otherwise, I don't understand why AI companies keep using BPE for their token sets. Unigram methods leads to more legible tokens, fewer glitch tokens, fewer super-long outlier tokens, etc.