9 ms·
This is too metaphorical, but, still, basically correct. Nice to see that. Essentially, in the latent / embedding / semantic space, "seahorse emoji" is somethi
by D-Machine 1y ago
This is too metaphorical, but, still, basically correct. Nice to see that.
Essentially, in the latent / embedding / semantic space, "seahorse emoji" is something that is highly probable. Actually, more accurately, since LLMs aren't actually statistical or probabilistic in any serious sense, "seahorse emoji", after tokenization and embedding, is very close to the learned manifold, and other semantic embeddings involving related emoji are very close to this "seahorse emoji" tokenization embedding.
An LLM has to work from this "seahorse emoji" tokenization embedding position, but can only make outputs through the tokenizer, which can't accurately encode "seahorse emoji" in the first place. So, you get a bunch of outputs that are semantically closest to (but still far from) a (theoretical) seahorse emoji. Then, on recursive application, since these outputs are now far enough from the the sort of root / foundational position on the manifold, the algorithm probably is doing something like an equivalent of a random walk on the manifold, staying close to wherever "seahorse emoji" landed, but never really converging, because the tokenization ensures that you can never really land back "close enough" to the base position.
I.e. IMO this is not as much a problem with (fixed) tokenization of the inputs, but moreso that tokenization of the outputs is fixed.
- mh- 1y agoThis explanation was very understandable, thank you for taking the time to write it.
- bravura 1y agoYou're missing one key point, which is what makes this failure mode unusual. Namely, that there is (incorrect) knowledge in the training data that "seahorse emoji" exists. So when prompted: "Does [thing you strongly believe exist]?" the LLM must answer: "Yes, ..." (The second nuance is that the LLM is strongly encouraged to explain its answers so it receives a lower score just by saying only "Yes.") But I and probably others appreciate your more detailed description of how it enters a repair loop, thank you. [edit: I disagree that LLMs are not statistical or probabilistic, but I'm not sure this is worth discussing.] [edit 2: Google is no longer telling me how many web pages a term responds, but "seahorse emoji" and "lime emoji" quoted both return over ten pages of results. The point being that those are both 'likely' terms for an LLM, but only the former is a likely continuation of 'Does X exist? Yes, ..."]
- D-Machine 1y agoYou're right, seahorse emoji is almost certainly in the training data, so we should amend my explanation to say that "seahorse emoji" is not just close to the training manifold, but almost certainly right smack on it. The rest of what I said would still apply, and my explanation would also to apply to where other commenters note that this behaviour is emitted to some degree with similar other "plausible" but non-existent emoji (but which are less likely to be in the training data, a priori). EDIT FOR THIS PARAGRAPH ONLY: Technically, on reflection, since all fitting methods employ regularization methods, it is still in fact unlikely the fitted manifold passes exactly through all / most training data points, and saying that "seahorse emoji" is "very close" to the training manifold is still actually technically probably most accurate here. You're also right that it is a long discussion to say to what extent LLMs are statistical or probabilistic, but, I would maybe briefly say that if one looks into issues like calibration, conformal prediction, and Bayesian neural nets, it is clear most LLMs that people are talking about today are not really statistical in any serious sense (softmax values are scores, not probabilities, and nothing about pre-training or tuning typically involves calibration—or even estimation—in LLMs). Yes, you can use statistics to (help) explain the behaviour of deep models or certain layers (usually making assumptions that are of dubious relevance to actual practice), but geometric analogies, regularization methods, and matrix conditioning intuitions are what have clearly guided almost all major deep learning advances, with statistical language and theory largely being post-hoc, hand-wavey, and (IMO) for the purpose of publication / marketing. I really think we could de-mystify a huge amount of deep learning if we were just honest it was mostly fancy curve fitting with some intuitive tricks for smoothing and regularization that clearly worked long before any rigorous statistical justification (or which still clearly work in complicated ways, despite such an absence of statistical understanding; e.g. dropout, norm layers, the attention layer itself, and etc). Just, it gets complicated when you get into diffusion models and certain other specific models that are in fact more explicitly driven by e.g. stochastic differential equations and the like.
- bravura 1y ago"my explanation would also to apply to where other commenters note that this behaviour is emitted to some degree with similar other "plausible" but non-existent emoji (but which are less likely to be in the training data, a priori)." I agree with you partially. I just want to argue there are several factors that lead to this perverse behavior. Empirically: Use web gpt-5-instant in TEMPORARY mode. If you ask for "igloo emoji" it confidently (but ONLY in temporary mode) says that "Yes, igloo emoji is in Unicode 12 and is [house-emoji ice-emoji]." Then it basically stops. But it has satisfied its condition of confidently expressing its false knowledge. (Igloo emoji doesn't exist. gpt-5-instant in non-temporary mode says no. This is also weird because it suggests the temporary mode system prompt is laxer or different.) The mechanism you describe partially explains why "seahorse emoji" leads to babbling: As it outputs the next token, it realizes that the explanation would be worse off it if next emits stop token, so instead it apologizes and attempts to repair. And cannot satisfy its condition of expressing something confidently. The upstream failure is poor knowledge. That combined with being tuned to be helpful and explanatory, and having no grounding (e.g. websearch) forces it to continue. Finally, the token distance from the manifold is the final piece of the puzzle in this unholy pathological brew. You're incorrect that statistical language modeling is "post-hoc", it's rather "pre-hoc" / "pre-hack". Most foundational works in language modeling started as pure statistical models (for example, classic ngram models and Bengio's original neural language model from 2003), and it was later that hacks got introduced that removed statistical properties but actually just worked (Collobert and Weston 2008, as influenced by Bottou and LeCun). Where I agree with you is that we should have done away with the statistical story long ago. LeCun's been on about energy-based models forever. Even on HN last week, punters criticize him that JEPA hasn't had impact yet, as if he were behind the curve instead of way ahead of it. People like statistical stories but, similarly to you, I also think they are a distraction.
- kqr 1y agoBut wait, if the problem is the final tokenisation, what would happen if we stopped it one or two layers before the final layer? I get that the result would not be as readable to a human as the final layer, but would it not be as confused with its own output anymore? Or would it still be a problem because we're collapsing a distribution of likely responses down to a single response, and it's not happy with that single response even if it is fuzzier than what comes out of the last layer?
- D-Machine 1y agoIt's not so clear how one could use the output of an embedding layer recursively, so it is a bit ill-defined to know what you mean by "stopped it" and "confused with its own output" here. You are mixing metaphor and math, so your question ends up being unclear. Yes, the outputs from a layer one or two layers before the final layer would be a continuous embedding of sorts, and not as lossy (compared to the discretized tokenization) at representing the meaning of the input sequence. But you can't "stop" here in a recursive LLM in any practical sense.