2 ms·
> I feel like the open source LLMs that are out there are never advertised/compared by their vocab list size. Increasing the vocab size increases training cost
by filterfiber 3y ago
> I feel like the open source LLMs that are out there are never advertised/compared by their vocab list size.
Increasing the vocab size increases training costs with little improvement in evaluation performance (how "smart" it is) and relatively not a ton of evaluation speed improvement. Words not in the dictionary can be made with several tokens. GPT3.5turbo/GPT4 uses 2 for "compared", "comp" and "ared".
That's not to say the existing vocab lists can't be further optimized, but there's been a lot more focus on the parameter count, structure, training/finetuning optimizations like LoRA, quantization methods, and training data, as these are what actually embed the information of how to predict the correct token.
There are a few cases where the vocab list is very important and you will see it mentioned. The more human languages you want to support the more tokens you'll generally want. GPT3's old tokenizer used to not have 4 spaces " " as a token which wasn't great for programming, so their "codex" model had a different tokenizer that did.
Other cases for specialized tokens include special characters like "fill-in-the-middle", control tokens for "<System>","</System>","<Prompt>","</Prompt>", and a few things of that nature.
A lot of llama models have added a few extra control tokens for a few different purposes. Note that because tokens are mapped from a number - you can generally add a new token and just finetune the model a bit to embed the information in the model, but you generally don't want to change existing tokens which have already had their usage "baked" into the model.