3 ms·
This thing will still hallucinate, not matter what new bells and whistles have been attached to it, meaning it will never be used for anything important and cri
by optimalsolver 2y ago
This thing will still hallucinate, not matter what new bells and whistles have been attached to it, meaning it will never be used for anything important and critical in the real world.
- barfbagginus 2y agoHere's a system that uses an llm to generate equivalence proofs for refactoring operations. https://news.ycombinator.com/item?id=40634775 https://news.ycombinator.com/item?id=40634775 In this system, the llm can hallucinate to its hearts content - the hallucinations are then fed into a proof engine and if they are a valid proof then it wasn't a hallucination, and the computation succeeds. If it fails it just tries again. So hallucinations cannot actually leave the system, and all we get are valid refactorings with working proofs of validity. Binding the LLM to a formal logic and proof engine is one way to stop them hallucinating and make them useful for the real world. But you would have to actually care about Proof and Truth to concede any point here. If you're only protecting the worldview where AI can never do things that humans can do, then you're going to have to retreat into some form of denial. But if you are interested in actual ways forward to useful AI, then results like this should give you some hope! Good luck and good day either way!
- freilanzer 2y ago> Binding the LLM to a formal logic and proof engine is one way to stop them hallucinating and make them useful for the real world. Checking the output does not mean the model does not hallucinate and thus does not help for all other cases in which there is no "formal logic and proof engine".
- barfbagginus 2y agoWhat if I consider the model to be the llm plus whatever extra components it has that allows it to not hallucinate? In that case then the model doesn't hallucinate, because the model is the llm plus the bolt-ons. Remember llm truthers claim that no bolt-ons can ever fully mitigate an llm's hallucinations. And yet in this case it does. But saying that it doesn't matter because other llms will still hallucinate is moving the goal post, or at least discounting the utility of this incremental progress. I think it's unfair to do this because there are many many domains where things can indeed be reduced to a formal logic amenable to approve engine. If they don't care about the actual output of a hybrid system that doesn't hallucinate, because it's math and not speech, then do they care about solving the issue at all, or providing human utility? I get the feeling that they only want to be right, not the benefit anyone. This shows that in cases where we can build good enough verifiers, hallucinations in a component of the system do not have to poison the entire system. Our own brains work this way - we have sections of our brains that hallucinate, and sections of the brain that verify. When the sections of the brain that verify our sleep, we end up hallucinating dreams. When the sections of the brain that verify are sick, we end up hallucinating while awake. I agree with you that the current system does not solve the problem for natural language. However it gives an example of a non hallucinating hybrid llm system. So the problem is reduced from having to make llms not hallucinate at all, to designing some other systems, potentially not an llm at all, that can reduce the number of hallucinations to a useful amount.
- barfbagginus 2y agoYou have no proof that every modification of the architecture will continue to have hallucinations. How could you prove that? Even LeCunn admits that the right modification could solve the issue. You're trying to make this point in a circular way - saying it's impossible just because you say it's impossible - for some reason other than trying to get to the bottom of the truth. You want to believe that there's some kind of guarantee that no offspring of the auto regressive architecture can never get rid of hallucinations. I'm saying they're simply no such guarantee.
- barfbagginus 2y agoPlus, humans bullshit all the time, even well paid and highly trained humans like Doctors and Lawyers. They will bullshit while charging you 400 an hour. Then they'll gaslight you if you try to correct their bullshit. AI will bullshit sometimes, but you can generally call it on the bullshit and correct it. For the tasks that it helps me with, I could work with a human. But the human I could afford would be a junior programmer. Not only do they bullshit more than a well prompted AI, but I also have to pay them 30 an hour, and they can't properly write specs or analyze requirements. GPT 4 can analyze requirements. Much better then a junior, and I'm many ways better than me. For pennies. I do use it in the real world, to maintain and develope software for the industrial design company I own. It would be foolish if I didn't. I've been able to modernize all our legacy codes and build features that used to stump me. Maybe the fact is that I'm an incompetent programmer, and that's why I find it so helpful. If that's the case so be it! It's still a significant help that is accessible to me. That matters!