4 ms·
> are LLMs just a large way of spitting out which token it thinks is best based on scoring? aka it's guessing at best? They are 100% exactly this. More specifi
by filterfiber 3y ago
> are LLMs just a large way of spitting out which token it thinks is best based on scoring? aka it's guessing at best?
They are 100% exactly this. More specifically they spit out a list of "logits" aka the probability of every token (llama has a vocab size of 32000, so you'll get 32000 probabilities).
Taking the highest probability is called "greedy sampling". Often you actually want to sample from either the top 5 (top k sampling), or the ones over say 90% (top p).
If you're doing things like programming or doing multiple choice questions, you can choose to only sample from those logits - for example, if the only outputs should be "true" or "false", ignore any token that isn't "t", "tr", "true", "f", "fa", "false", etc. This makes it adhere to a schema.
> How will they ever grow to the point where they are nothing more than just really good at guessing what to regurgitate based on what it's already seen?
This is what humans do the big difference is our "weights" aren't frozen and we update them in realtime based on real world feedback. Once you loop in the real world feedback you get much better results, for example if a llm recommends a command with a syntax error, if you give it back the error message it can correct it. This is why training on synthetic data is possible.
> will they ever be trustable/perfect?
This is why I hate the word "hallucinate", it's not a hallucination, it's a "miss-prediction". Humans do the exact same thing all the time - miss remembering, misspeaking, doing fast math, etc.
In many ways LLMs are already more "trustable" than inexperienced humans. The main thing you need is some mechanism of double-checking the output, the same as we do with humans. We can "trust" them once their probability of acceptable answers is high enough.
- MuffinFlavored 3y ago> More specifically they spit out a list of "logits" aka the probability of every token (llama has a vocab size of 32000, so you'll get 32000 probabilities). I would have thought you have 7B probabilities (7b possible tokens) and 32k is just the context. So for every token (up to 32k, because that's when it runs out of context/size), you have a 7b probability. I feel like you mixed up context size with # of parameters?
- filterfiber 3y ago> I feel like you mixed up context size with # of parameters? Respectfully I do not think I did :) The 7B is the parameter count. 32k is the vocab size, or list of tokens. I don't have access to the llama repo at the moment so I'll use this one - https://huggingface.co/NousResearch/Llama-2-7b-chat-hf/blob/main/config.json#L24 https://huggingface.co/NousResearch/Llama-2-7b-chat-hf/blob/... The vocab list is here https://huggingface.co/NousResearch/Llama-2-7b-chat-hf/raw/main/tokenizer.json https://huggingface.co/NousResearch/Llama-2-7b-chat-hf/raw/m... vocab_size is 32k. max_position_embeddings is 4096 for the context. Note that you can use a longer or shorter context length, but unless you use some tricks, longer context will result in rapidly decreasing performance. You will get 32k logits as the end result.
- MuffinFlavored 3y agoCan you help me understand how a parameter and a token differ? > 32k is the vocab size, or list of tokens. This sounds to me like it is choosing from 1 of 32k tokens when it is scoring/generating an answer. Where does the 7,000,000,000 parameters come from then? I would have thought it is picking from 1 of 7B parameters? Parameter != token?
- filterfiber 3y ago> Parameter != token? They are two different things. A parameter is not what programmers call "parameters/arguments", it's the total number of weights from all of the layers in the network. A token is a number that represents some character(s), there is a map of them in the vocab list. Similar to ascii or UTF8, but a token can be either a single character or multiple. For example "I" might be a token and "ca" might be a token. Shorter words especially might be a single token, but most words are a combination of tokens. > This sounds to me like it is choosing from 1 of 32k tokens when it is scoring/generating an answer. [...] I would have thought it is picking from 1 of 7B parameters? The model doesn't do that itself. The last layer of the model maps it to a single list that is 32k long, and outputs that of tokens and each ones probability (the combination of a (token, probability) is called a logit). So the output is a list of logits. The step after this is external to the model - sampling. You can choose the highest scoring token if you want (greedy sampling), but usually you want to make it a little random with either one of the top 20 tokens (top_k) or from the best 90th percentile tokens (top_p). There's other fun tricks like logit biasing. Finally you map the tokens to their respective values from the vocab list in my previous comment. > Where does the 7,000,000,000 parameters come from then? Here's a walk through for how to count the parameters for Llama-13B. It's worth noting that most of these models are commonly rounded (that's why you might see llama-33B be called llama-30B). https://medium.com/@saratbhargava/mastering-llama-math-part-1-a-step-by-step-guide-to-counting-parameters-in-llama-2-b3d73bc3ae31 https://medium.com/@saratbhargava/mastering-llama-math-part-... https://github.com/saratbhargava/ai-blog-resources/blob/main/LLM/Llama_2_param_count.ipynb https://github.com/saratbhargava/ai-blog-resources/blob/main...