8 ms·
Spoiler: You're never going to get rid of hallucinations in the autoregressive, next token prediction paradigm (aka LeCun's Law). The issue here is people tryi
by optimalsolver 2y ago
Spoiler: You're never going to get rid of hallucinations in the autoregressive, next token prediction paradigm (aka LeCun's Law).
The issue here is people trying to use language models as deterministic problem solvers, rather than for what they actually excel at (semi-creative text generation).
- plewd 2y agoIs LeCun's Law even a thing? Searching up for it doesn't yield many results, except for a HN comment where it has a different definition. I guess it could be from some obscure paper, but with how poorly it's documented it seems weird to bring it up in this context.
- vjerancrnjak 2y ago“Label bias” or “observation bias” a phenomenon where going outside of the learned path lives little room for error correction. Lecun talks about the lack of joint learning in LLMs.
- mdp2021 2y agoA reference could be this: https://futurist.com/2023/02/13/metas-yann-lecun-thoughts-large-language-models-llms/ https://futurist.com/2023/02/13/metas-yann-lecun-thoughts-la... (Speaking of "law" is rhetoric, but an idea is pretty clear.)
- YeGoblynQueenne 2y agoI think the OP may be referring to this slide that Yann LeCun has presented on several occasions: https://youtu.be/MiqLoAZFRSE?si=tIQ_ya2tiMCymiAh&t=901 https://youtu.be/MiqLoAZFRSE?si=tIQ_ya2tiMCymiAh&t=901 To quote from the slide: * Probability e that any produced token takes us outside the set of correct answers * Probability that answer of length n is correct * P(correct) = (1-e)^n * This diverges exponentially * It's not fixable (without a major redesign)
- atq2119 2y agoDoesn't that argument make the fundamentally incorrect assumption that the space of produced output sequence has pockets where all output sequence with a certain prefix are incorrect? Design your output space in such way that every prefix has a correct completion and this simplistic argument no longer applies. Humans do this in practice by saying "hold on, I was wrong, here's what's right". Of course, there's still a question of whether you can get the probability mass of correct outputs large enough.
- marcosdumay 2y agoHow do you do this in something where the only memory is the last few things it said or heard?
- roboboffin 2y agoIs this similar to the effect that I have seen when you have two different LLMs talking to each other, they tend to descend into nonsense ? A single error in one of the LLM's output and that then pushes the other LLM out of distribution. I kind of oscillatory effect when the train of tokens move further and further out of the distribution of correct tokens.
- diggan 2y ago> Is this similar to the effect that I have seen when you have two different LLMs talking to each other, they tend to descend into nonsense ? Is that really true? I'd expect that with high temperature values, but otherwise I don't see why this would happen, and I've experimented with pitting same models against each other and also different models against different models, but haven't come across that particular problem.
- reportgunner 2y agoCan you show examples ? In any AI related discussions there are only some claims by people and never examples of the AI working well.
- whimsicalism 2y agoIt’s a thing in that he said it but it’s not an actual law and it has several obvious logical flaws. It applies just as equally to human utterances.
- famouswaffles 2y agoIf you're talking about label bias then you don't need to solve label bias to 'solve' hallucinations when the model has already learnt internally when it's bullshitting or going off the rails.
- seydor 2y ago"never" is not itself a problem, people do the same you only need to solve fusion correctly once
- shawnz 2y agoDoes anyone here know, has anyone tried something like feeding the perplexity of previous tokens back into the model, so that it has a way of knowing when it's going off the rails? Maybe it could be trained to start responding less confidently in those cases, reducing its desire to hallucinate.
- famouswaffles 2y agoModels already know when they are going off the rails. https://news.ycombinator.com/item?id=41504226 https://news.ycombinator.com/item?id=41504226. That's not the problem. The problem is that they don't care to tell you.
- whimsicalism 2y agoLeCuns argument is seriously flawed. It is not at all a rigorous one and you should not make such sweeping statements based on nothing.
- barbarr 2y agoAt this point I just invert everything LeCun says about AI. Chances are he'll flip flop on his own statement a few months later anyways.
- whiplash451 2y agoLeCun has been pretty steady for years now.
- wpietri 2y agoVery nice to see this point being made. One way I explain it to people: Imagine a corporation that only has a PR department. Extremely good at generating press releases and answering reporter questions. But without the rest of the company, the output text isn't constrained by anything meaningful. In an alternate universe, one where people understood this, people would be using LLMs for nothing serious, but a whole lot of fun little art projects.