4 ms·
The problem can be mitigated by using bigger models or training embeddings, but it cannot be eliminated. At the end of the day, it remains a sequence completio
by usrbinbash 3y ago
The problem can be mitigated by using bigger models or training embeddings, but it cannot be eliminated.
At the end of the day, it remains a sequence completion engine. An LLM can only care about the statistical likelyhood of sequences, that the entire universe it exists in.
And if a sequence making a wrong statement happens to be more likely than, or close enough for the set "heat" of the algorithm to, a sequence denoting a correct statement, then there is a chance the model will hallucinate.
If it were otherwise, we would already have a good solution for prompt injection attacks, and we don't.
- namaria 3y agoLarge language models do a great job of emulating meaning. That's their biggest strength and greatest weakness. We got a finally good chatbot, at the cost of making people project so much more then a chatbot on it. LLMs can't mean anything. They can only produce surrogate meaning. When you compress vast amounts of text into a model and query the model you can only get back information that is already in the training data set. It's inherently backwards looking, it cannot produce lowered entropy forward in time. You can't get more knowledge out of it, only less then you inputted in the first place.
- cookieperson 3y agoAlso, it's important to state that, the BEST case scenario is that you get the knowledge out that went in. The common scenario would be that you have loss. It's also important to remember this thing was trained on Reddit, etc. We all know how much complete trash, bot posts, hot takes, and conflicting views there are on there.
- namaria 3y agoGood point. It wasn't even trained on high grade knowledge, just on very noise bodies of language.