3 ms·
> Is the "parameters" the vocab list? The vocab list is the mapping of tokens to their utf encodings. for example, the token 680 is "ide", token 1049 is "land
by filterfiber 3y ago
> Is the "parameters" the vocab list?
The vocab list is the mapping of tokens to their utf encodings. for example, the token 680 is "ide", token 1049 is "land".
> So from the 7B parameters, it picks which 32k potential parameters to use as tokens? So the "vocab list" is a 32k subset of the 7B parameters that changes with each request?
No, do not think of parameters and tokens as comparable in any way (at least right now). The parameter count is just the total number of weights and biases throughout all of the layers. The vocab size is static, and will never change during inference.
this is simplified - you give it a list of tokens, and then there's a bunch of linear algebra with matrices of "weights". All of those weights combined is the 7B parameters.
The final operation results in a 32k long list of probabilities. The 680'th item in that list is the probability of "ide", the 1049th item is the probability of "land".
The model never "picks" anything, there are no "if" statements so to speak, you give it an input and it does a bunch of multiplication resulting in a list of predictions 32k long.
The model does not "pick the best token" at any point. It simply hands you a 32k predictions. It's your job to pick the token, you can certainly just pick the highest probability but that's not usually the sampling method people use https://towardsdatascience.com/how-to-sample-from-language-models-682bceb97277 https://towardsdatascience.com/how-to-sample-from-language-m...
I highly recommend 3blue1brown to learn more about neural networks (and anything math related he's great) https://www.3blue1brown.com/lessons/neural-networks https://www.3blue1brown.com/lessons/neural-networks
tl;dr - very oversimplified - you multiply/add 7B numbers with your input list resulting in 32k predictions, which you map to the letters in the vocab list.
- MuffinFlavored 3y ago> The vocab size is static, and will never change during inference. I feel like the open source LLMs that are out there are never advertised/compared by their vocab list size. 7B weights to pick which of 32k tokens to pick, over and over (per token, sequentially)
- filterfiber 3y ago> I feel like the open source LLMs that are out there are never advertised/compared by their vocab list size. Increasing the vocab size increases training costs with little improvement in evaluation performance (how "smart" it is) and relatively not a ton of evaluation speed improvement. Words not in the dictionary can be made with several tokens. GPT3.5turbo/GPT4 uses 2 for "compared", "comp" and "ared". That's not to say the existing vocab lists can't be further optimized, but there's been a lot more focus on the parameter count, structure, training/finetuning optimizations like LoRA, quantization methods, and training data, as these are what actually embed the information of how to predict the correct token. There are a few cases where the vocab list is very important and you will see it mentioned. The more human languages you want to support the more tokens you'll generally want. GPT3's old tokenizer used to not have 4 spaces " " as a token which wasn't great for programming, so their "codex" model had a different tokenizer that did. Other cases for specialized tokens include special characters like "fill-in-the-middle", control tokens for "<System>","</System>","<Prompt>","</Prompt>", and a few things of that nature. A lot of llama models have added a few extra control tokens for a few different purposes. Note that because tokens are mapped from a number - you can generally add a new token and just finetune the model a bit to embed the information in the model, but you generally don't want to change existing tokens which have already had their usage "baked" into the model.