5 ms·
Yann LeCun would say [0] that it's autoregression itself that's the problem, and ML of this type will never bring us anywhere near AGI. At the very least you c
by optimalsolver 2y ago
Yann LeCun would say [0] that it's autoregression itself that's the problem, and ML of this type will never bring us anywhere near AGI.
At the very least you can't solve the hallucination problem while still in the autoregression paradigm.
[0] https://twitter.com/ylecun/status/1640122342570336267 https://twitter.com/ylecun/status/1640122342570336267
- vessenes 2y agoI think this method might not be amenable to the exponential divergence argument actually. Depending on token sampling methods, this one could look at a proposed generation as a whole and revise it. I’m not sure the current token sampling method they propose does this right now, but I think it’s possible with the information they get out of the probabilities.
- modeless 2y agoYes, to me this seems to address LeCun's objection, or at least point the way to something that does. It seems possible to modify this into something that can identify and correct its own mistakes during the sampling process.
- vessenes 2y agoWell, I think I understand LeCun has a broader critique that any sort of generated-in-a-vacuum text which doesn't interact with meatspace is fundamentally going to be prone toward divergence. Which, I might agree with, but is also, just, like, his opinion, man. Or put less colloquially, that's a philosophical stance sitting next to the math argument for divergence. I do think this setup can answer (much of) the math argument.
- _w1tm 2y agoDoes everything have to take us towards AGI? If someone makes a LLM that’s faster (cheaper) to run then that has value. I don’t think we want AGI for most tasks unless the intent is to produce suffering in sentient beings.
- ben_w 2y ago> I don’t think we want AGI for most tasks unless the intent is to produce suffering in sentient beings. Each letter of "AGI" means different things to different people, and some use the combination to mean something not present in any of the initials. The definition OpenAI uses is for economic impact, so for them, they do want what they call AGI for most tasks. I have the opposite problem with the definition, as for me, InstructGPT met my long-standing definition of "artificial intelligence" while suddenly demonstrating generality in that it could perform arbitrary tasks rather than just next-token prediction… but nobody else seems to like that, and I'm a linguistic descriptivist, so I have to accept words aren't being used the way I expected and adapt rather than huff.
- barfbagginus 2y agoI call GPT an AGI 1. To highlight that the system passes the turing test and has general intelligence abilities beyond the median human 2. To piss off people who want AGI to be a God or universal replacement for any human worker or intellectual The problem with AGI as a universal worker replacement - the way that it can lead to sentient suffering - is the presumption that these universal worker replacements should be owned by automated corporations and hyper wealthy individuals, rather than by the currently suffering sentient individuals who actually need the AI assistance. If we cannot make Universal Basic AGI that feeds and clothes everyone by default as part of the shared human legacy - UBAGI - then AGI will cause harm and suffering.
- ben_w 2y ago> 1. To highlight that the system passes the turing test and has general intelligence abilities beyond the median human I think that heavily depends on what you mean by "intelligence", which in turn depends on how you want to make use of it. I would agree that it's close enough to the Turing test as to make the formal test irrelevant. AI training currently requires far more examples than any organic life. It can partially make up for this by transistors operating faster than synapses by the same ratio to which marathon runners are faster than continental drift — but only partially. In areas where there is a lot of data, the AI does well; in areas where there isn't, it doesn't. For this reason, I would characterise them as what you might expect from a shrew that was made immortal and forced to spend 50,000 years reading the internet — it's still a shrew, just with a lot of experience. Book smarts, but not high IQ. With LLMs, the breadth of knowledge makes it difficult to discern the degree to which they have constructed a generalised world model vs. have learned a lot of catch-phrases which are pretty close to the right answer. Asking them to play chess can result in them attempting illegal moves, for example, but even then they clearly had to build a model of a chess board good enough to support the error instead of making an infinitely tall chess board in ASCII art or switching to the style of a chess journalist explaining some famous move. For a non-LLM example of where the data-threshold is, remember that Tesla still doesn't have a level 4 self-driving system despite millions of vehicles and most of those operating for over a year. If they were as data-efficient as us, they'd have passed the best human drivers long ago. As is, while they have faster reactions than we do and while their learning experiences can be rolled out fleet-wide overnight, they're still simply not operating at our level and do make weird mistakes.
- cs702 2y agoLeCun may or may not be right, but I'm not sure this is relevant to the discussion here. The OP's authors make no claims about how their work might help get us closer to AGI. They simply enable autoregressive LLMs to do new things that were not possible before.
- TheEzEzz 2y agoLeCun is very simply wrong in his argument here. His proof requires that all decoded tokens are conditionally independent, or at least that the chance of a wrong next token is independent. This is not the case. Intuitively, some tokens are harder than others. There may be "crux" tokens in an output, after which the remaining tokens are substantially easier. It's also possible to recover from an incorrect token auto-regressively, by outputting tokens like "actually no..."
- sebzim4500 2y agoLeCun is a very smart guy but his track record predicting limitations of autoregressive LLMs is terrible.
- barfbagginus 2y agoCan I please convert you into someone who summarily barks at people for making the LeCunn Fallacy rather than making the LeCunn Fallacy yourself? And can you stop talking about AGI when it's not relevant to a conversation? Let's call that the AGI fallacy - the argument that a given development is worthless - despite actual technical improvements - because it's not AGI or supposedly can't lead to AGI. It's a problem. Every single paper on transformers has some low information comment to the effect of, "yeah, but this won't give us AGI because of the curse of LeCunn". The people making these comments never care about the actual improvement, and are never looking for improvements themselves. It becomes tiring to people, like yours truly :3, who do care about the work. Let's look at the structure of the fallacy. You're sidestepping the "without a major redesign" in his quote. That turns his statement from a statement of impossibility into a much weaker statement saying that auto regressive models currently have a weakness. A weakness which could possibly be fixed by redesign, which LeCunn admits. In fact this paper is a major redesign. It solves a parallelism problem, rather than the hallucination problem. But it still proves that major redesigns do sometimes solve major problems in the model. There could easily arise a regressive model that allows progressive online updating from an external world model - that's all it takes to break LeCunn's curse. There's no reason to think the curse can't be broken by redesign.
- optimalsolver 2y agoThis thing will still hallucinate, not matter what new bells and whistles have been attached to it, meaning it will never be used for anything important and critical in the real world.
- barfbagginus 2y agoHere's a system that uses an llm to generate equivalence proofs for refactoring operations. https://news.ycombinator.com/item?id=40634775 https://news.ycombinator.com/item?id=40634775 In this system, the llm can hallucinate to its hearts content - the hallucinations are then fed into a proof engine and if they are a valid proof then it wasn't a hallucination, and the computation succeeds. If it fails it just tries again. So hallucinations cannot actually leave the system, and all we get are valid refactorings with working proofs of validity. Binding the LLM to a formal logic and proof engine is one way to stop them hallucinating and make them useful for the real world. But you would have to actually care about Proof and Truth to concede any point here. If you're only protecting the worldview where AI can never do things that humans can do, then you're going to have to retreat into some form of denial. But if you are interested in actual ways forward to useful AI, then results like this should give you some hope! Good luck and good day either way!