3 ms·
I was saying that the seahorse emoji failure is not a tokenization issue. If you ask an LLM to do research, you will sometimes get hallucinated articles -- pote
by jsrozner 1y ago
I was saying that the seahorse emoji failure is not a tokenization issue. If you ask an LLM to do research, you will sometimes get hallucinated articles -- potentially plausible articles that, if they existed, would have been embedded at the position from which the model tried to decode. This is what we see happening with the seahorse emoji. The model identifies where the seahorse emoji would have been embedded if it existed and then decodes from that position.
In the research case you get articles that were never written. In the seahorse case later layers hallucinate the seahorse emoji, but in the final decoding step, output gets mapped onto another nearby emoji.
Admittedly, in one way the seahorse example is different from the research case. Article titles, since they use normal characters, can be produced whether they exist or not (e.g., "This is a fake hallucinated article" gets produced just as easily as "A real article title"). It's actually nice that the model can't produce the seahorse emoji since it gets forced (by tokens, yes) to decode back into reality.
Yes, tokenization affects how the hallucination manifests, but the underlying problem is not a tokenization one.