8 ms·
Entropy of a Large Language Model output
- WhitneyLand 2y ago> the output token of the LLM (black box) is not deterministic. Rather, it is a probability distribution over all the available tokens How is this not deterministic? Randomness is intentionally added via temperature.
- cjtrowbridge 2y agoEntropy is also added via a random seed. The model is only deterministic if you use the same random seed.
- HarHarVeryFunny 2y agoI think you're confusing training and inference. During training there are things like initialization, data shuffling and dropout that depend on random numbers. At inference time these don't apply.
- jampekka 2y agoDecoding (sampling) uses (pseudo) random numbers. Otherwise same prompt would always give the same response. Computing entropy generally does not. See e.g. https://huggingface.co/blog/how-to-generate https://huggingface.co/blog/how-to-generate
- HarHarVeryFunny 2y agoSure - but that's not the output of the model itself, that's the process of (typically) randomly sampling from the output of the model.
- throwaway314155 2y agoRight, sampling from a model, also known as *inference* (for LLM's). The inference here is perhaps less pure than what you refer to but you're talking to human beings; there's no need for heavy pedantry.
- hansvm 2y agoThe output "token" Yes, you can sample deterministically, but that's some combination of computationally intractable and only useful on a small subset of problems. The black box outputting a non-deterministic token is a close enough approximation for most people.
- HarHarVeryFunny 2y agoThe author of the article seems confused, saying: "The important thing to remember is that the output token of the LLM (black box) is not deterministic. Rather, it is a probability distribution over all the available tokens in the vocabulary." He is saying that there is non-determinism in the output of the LLM (i.e. in these probability distributions), when in fact the randomness only comes from choosing to use a random number generator to sample from this output.
- fancyfredbot 2y agoThe author is saying that the output token is not deterministic. I don't think they said the distribution was stochastic. Even so the distribution of the second token output by the model would be stochastic (unless you condition on the first token). So in that sense there may also be a stochastic probability distribution.
- hansvm 2y agoMostly unrelated (I agree with you, and I'm some ancestory comment you're responding to with the same line of thinking), I have built a couple LLMs where the distribution itself is stochastic. That's not key to how they work as a black box, but much like how quicksort has certain performance characteristics I did find it advantageous to introduce randomness into the model itself. You could still easily model the next token as a conditional probability distribution though if you wanted; the computation of entropy just might be a bit spendier.
- apstroll 2y agoThe output distribution is deterministic, the output token is sampled from the output distribution, and is therefore not deterministic. Temperature modulates the output distribution, but sitting it to 0 (i.e. argmax sampling) is not the norm.
- Der_Einzige 2y agoRunning temperature of zero/greedy sampling (what you call "argmax sampling") is EXTREMELY common. LLMs are basically "deterministic" when using greedy sampling except for either MoE related shenanigans (what historically prevented determinism in ChatGPT) or due to floating point related issues (GPU related). In practice, LLMs are in fact basically "deterministic" except for the sampling/temperature stuff that we add at the very end.
- HarHarVeryFunny 2y ago> except for either MoE related shenanigans (what historically prevented determinism in ChatGPT) The original ChatCPT was based on GPT-3.5, which did not use MoE.
- alew1 2y ago"Temperature" doesn't make sense unless your model is predicting a distribution. You can't "temperature sample" a calculator, for instance. The output of the LLM is a predictive distribution over the next token; this is the formulation you will see in every paper on LLMs. It's true that you can do various things with that distribution other than sampling it: you can compute its entropy, you can find its mode (argmax), etc., but the type signature of the LLM itself is `prompt -> probability distribution over next tokens`.
- wyager 2y agoThe temperature in LLMs is a parameter of a regularization step that determines how neuron activation levels get mapped to odds ratios. Zero temperature => fully deterministic The neuron activation levels do not inherently form or represent a probability distribution. That's something we've slapped on after the fact
- alew1 2y agoAny interpretation (including interpreting the inputs to the neural net as a "prompt") is "slapped on" in some sense—at some level, it's all just numbers being added, multiplied, and so on. But I wouldn't call the probabilistic interpretation "after the fact." The entire training procedure that generated the LM weights (the pre-training as well as the RLHF post-training) is formulated based on the understanding that the LM predicts p(x_t | x_1, ..., x_{t-1}). For example, pretraining maximizes the log probability of the training data, and RLHF typically maximizes an objective that combines "expected reward [under the LLM's output probability distribution]" with "KL divergence between the pretraining distribution and the RLHF'd distribution" (a probabilistic quantity).
- apstroll 2y agoUnder a crossentropy loss the output activations do absolutely represent a probability distribution, since that is what we're modeling.
- TeMPOraL 2y agoThere's extra randomness added accidentally in practice: inference is a massively parallelized set of matrix multiplications, and floating point math is not commutative - the randomness in execution order gets converted into a random FP error, so even setting temperature to 0 doesn't guarantee repeatable results.
- HeatrayEnjoyer 2y agoOnly if the inference software doesn't guarantee concurrency, which is CS 101
- pizza 2y agoThis sort of nondeterministic scheduling of non associative floating point ops is essentially running at the level of GPU firmware, so, I would imagine that in this case, Nvidia is aware.
- nikkindev 2y agoAuthor here: Yes. You are right. I was meaning to paint a picture that instead of the next token appearing magically, it is sampled from a probability distribution. The notion of determinism could be explained differently. Thanks for pointing it out!
- netruk44 2y agoI wonder if we could combine ‘thinking’ models (which write thoughts out before replying) with a mechanism they can use to check their own entropy as they’re writing output. Maybe it could eventually learn when it needs to have a low entropy token (to produce a more-likely-to-be-factual statement) and then we can finally have models that actually definitely know when to say “Sorry, I don’t seem to have a good answer for you.”
- vletal 2y agohttps://github.com/xjdr-alt/entropix https://github.com/xjdr-alt/entropix
- Der_Einzige 2y agoEntropix will get it's time in the sun, but for now, the LLM academic community is still 2 years behind the open source community. Min_p sampling is going to end up getting an oral about it at ICLR with the scores it's getting... https://openreview.net/forum?id=FBkpCyujtS https://openreview.net/forum?id=FBkpCyujtS
- diggan 2y ago> the LLM academic community is still 2 years behind the open source community Huh, isn't it the other way around? Thanks to the academic (and open) research about LLMs, we have any open source community around LLMs in the first place.
- pizza 2y agoThere's a paper that probed how strongly a model would focus on prompt-supplied tokens when generating a response as a signal that it was trying to use the prompt as the source of information as opposed to knowledge it had been trained on. Ie, how much it was trying to lie based on it assuming that the information in the prompt was true, as opposed to having a rich internal model of the thing that is being verified. It looks like it works, sort of, sometimes, when you have access to the actual labels. The results from this work, in the more real-world unsupervised setting, are better than random, sure, but not good enough to really be exciting or reliable. https://arxiv.org/html/2402.03563v1 https://arxiv.org/html/2402.03563v1
- fedeb95 2y agoit seems very noise-like to me.
- gwern 2y agoYou are observing "flattened logits" https://arxiv.org/pdf/2303.08774#page=12&org=openai https://arxiv.org/pdf/2303.08774#page=12&org=openai . The entropy of ChatGPT (as well as all other generative models which have been 'tuned' using RLHF, instruction-tuning, DPO, etc) is so low because it is not predicting "most likely tokens" or doing compression. A LLM like ChatGPT has been turned into an RL agent which seeks to maximize reward by taking the optimal action. It is, ultimately, predicting what will manipulate the imaginary human rater into giving it a high reward. So the logits aren't telling you anything like 'what is the probability in a random sample of Internet text of the next token', but are closer to a Bellman value function, expressing the model's belief as to what would be the net reward from picking each possible BPE as an 'action' and then continuing to pick the optimal BPE after that (ie. following its policy until the episode terminates). Because there is usually 1 best action, it tries to put the largest value on that action, and assign very small values to the rest (no matter how plausible each of them might be if you were looking at random Internet text). This reduction in entropy is a standard RL effect as agents switch from exploration to exploitation: there is no benefit to taking anything less than the single best action, so you don't want to risk taking any others. This is also why completions are so boring and Boltzmann temperature stops mattering and more complex sampling strategies like best-of-N don't work so well: the greedy logit-maximizing removes information about interesting alternative strategies, so you wind up with massive redundancy and your net 'likelihood' also no longer tells you anything about the likelihood. And note that because there is now so much LLM text on the Internet, this feeds back into future LLMs too, which will have flattened logits simply because it is now quite likely that they are predicting outputs from LLMs which had flattened logits. (Plus, of course, data labelers like Scale can fail at quality control and their labelers cheat and just dump in ChatGPT answers to make money.) So you'll observe future 'base' models which have more flattened logits too... I've wondered if to recover true base model capabilities and get logits that actually meaningful predict or encode 'dark knowledge', rather than optimize for a lowest-common-denominator rater reward, you'll have to start dumping in random Internet text samples to get the model 'out of assistant mode'.
- cbzbc 2y agoSorry, which particular part of that paper are you linking to, the graph at the top of that page doesn't seem to link to your comment?
- EncomLab 2y agoWe should stop using the term "black box" to mean "we don't know" when really it's "we could find out but it would be really hard". We can precisely determine the exact state of any digital system and track that state as it changes. In something as large as a LLM doing so is extremely complex, but complexity does not equal unknowable. These systems are still just software, with pre-defined operations executing in order like any other piece of software. A CPU does not enter some mysterious woo "LLM black box" state that is somehow fundamentally different than running any other software, and it's these imprecise terms that lead to so much of the hype.
- Ecoste 2y agoSo going by your definition what would be a true black box?
- EncomLab 2y agoA starting point would be a system that does not require the use of a limited set of pre-defined operations to transform from one state to another state via the interpretation of a set of pre-existing instructions. This rules out any digital system entirely.
- achierius 2y agoBut what _would_ qualify? The point being made is that your definition is so constricting as to be useless. Nothing (sans perhaps true physical limit-conditions, like black-holes) would be a black box under your definition.
- EncomLab 2y agoIt's really only constricting to state machines which are dependent upon a fixed instruction set to function.
- saurik 2y agoThis is much more similar to the technique of obfuscating encryption algorithms for DRM schemes that I believe is often called "white-box cryptography".
- behnamoh 2y agoThis was discussed in my paper last year: https://arxiv.org/abs/2406.05587 https://arxiv.org/abs/2406.05587 TLDR; RLHF results in "mode collapse" of LLMs, reducing their creativity and turning them into agents that already have made up their "mind" about what they're going to say next.
- deleted 2y ago[deleted]
- nikkindev 2y agoAuthor here: Really interesting work. Updated original post to include link to the paper. Thanks!
- kleiba 2y agoIn LM research, it is more common to measure the exponentiation of the entropy, called perplexity. See also https://en.wikipedia.org/wiki/Perplexity https://en.wikipedia.org/wiki/Perplexity
- pona-a 2y agoPerhaps CoT and the like may be limited by this. If your model is cooked and does not adequately represent less immediately useful predictions, even if you slap a more global probability maximization mechanism, you can't extract knowledge that's been erased by RLHF/fine-tuning.
- K0balt 2y agoLow entropy is expected here, since the model is seeking a “best” answer based on reward training. But I see the same misconceptions as always around “hallucinations”. Incorrect output is just incorrect output. There is no difference in the function of the model, no malfunction. It is working exactly as it does for “correct “ answers. This is what makes the issue of incorrect output intractable. Some optimisation can be achieved through introspection, but ultimately, an llm can be wrong for the same reason that a person can be wrong, incorrect conclusions, bad data, insufficient data, or faulty logic/modeling. If there was a way to be always right, we wouldn’t need LLMs or second opinions. Agentic workflows and introspection/cot catch a lot, and flights of fancy are often not supported or replicated with modifications to context, because the fanciful answer isn’t reinforced in the training data. But we need to get rid of the unfortunate term for wrong conclusions,“hallucination” . When we say a person is hallucinating, it implies an altered state of mind. We don’t say that bob is hallucinating when he thinks that the sky is blue because it reflects the ocean, we just know he’s wrong because he doesn’t know about or forgot about Raleigh scattering. Using the term “hallucination” distracts from accurate thought and misleads people to draw erroneous conclusions.
- nikkindev 2y agoAuthor here: Wholeheartedly agree with your comment on hallucination. I initially set out to answer the question “Will entropy help identify hallucination?” And soon realised that it doesn’t, for the same reasons you mentioned above. So I pivoted to just writing about the entropy measure in the post. And this is also reflected by how I started with hallucination and then quickly veered away from it. I’ll be more careful in future posts & conversations. Thanks!
- K0balt 2y agoNice post, really, and I think it will help some people to understand more about how LLMs work, especially helping fix the dogma about “LLMs just randomly select the next most likely word” which is kinda true but so many qualifiers and contextual details apply that the statement is more misleading than useful. On undesired output, I would think it a great service to the field if we could come up with a better and earwormier word for “hallucinations” and somehow make it stick. Right now we have half the literate world walking around thinking that LLMs are licking frogs, and it does nothing to help people understand how to think about model outputs or how to increase the utility of these fantastic culture / data mining tools in their own lives.
- Lerc 2y agoThere is an interesting aspect of this behaviour used in the byte latent transformer model. Encoding tokens from source text can be done a number of ways, byte pair encoding, dictionaries etc. You can also just encode text into tokens (or directly into embeddings) with yet another model. The problem arises that if you are doing variable length tokens, how many characters do you put into any particular token, and then because that token must represent the text if you use it for decoding, where do you store count of characters stored in any particular token. The byte latent transformer model solves this by using the entropy for the next character. A small character model receives the history character by character and predicts the next one. If the entropy spikes from low to high they count that as a token boundary. Decoding the same characters from the latent one at a time produces the same sequence and deterministically spikes at the same point in the decoding indicating that it is at the end of the token without the length being required to be explicitly encoded. (disclaimer: My layman's view of it anyway, I may be completely wrong)
- deleted 2y ago[deleted]