4 ms·
Since the LLMs are probabalistic is it possible for the inference engines to highlight lower probability content. That might at least make them more "human" rat
by dougSF70 4y ago
Since the LLMs are probabalistic is it possible for the inference engines to highlight lower probability content. That might at least make them more "human" rather than being confident and inaccurate.
- nicognaw 4y agoSince LLMs generate words one by one, you can only determine whether the next word has a high probability. Some ideas occurred to me earlier, like applying prompt engineering to evaluate if the model is confident about the contents to be generated. Like "Please reply T if you are confident about the text that you are going to generate, and F if you are not". But if this could be done by architecting the neural networks, they would perform much better.
- pmoriarty 4y ago"Since LLMs generate words one by one, you can only determine whether the next word has a high probability" We don't fully understand how they work or what their limits are.
- RugnirViking 4y agowe do understand that they generate words with probabilities one by one though. They are transformers, the AI bit takes the entire prompt as input and returns a range of probabilities for the next token (token == word, it doesnt understand or see letters). A "dumb" algorithm then selects from these probabilities randomly using the probabilities returned by the ai based on the "temprature" setting (low temprature favors high probability words & high temprature selects from a whole lot of words but leads into really weird territory fast) It's also worth remembering that this is not just a markov chain as many people seem to think, it doesn't simply remember what words come next in its training set, because statistically if you take a random 10 consecutive words from a piece of text, chances are its never been written before. (also the trained model is much smaller than the size of the dataset, so to simply remember everything it would have to be the worlds best compression algorithm by orders of magnitude) Thats why we need an AI here, to learn the general rules of language so it can respond to chains of words it has never seen before. The sense that we "dont understand how they work" is that we dont know what the "rules of language" that it has learned are.
- pmoriarty 4y ago"we do understand that they generate words with probabilities one by one though" This is no more helpful in understanding AI's than is knowing that human brains operate according to the laws of physics is helpful in understanding the human mind.
- skybrian 4y agoYes, there are important mysteries that need to be solved through research [1] but there are still some limits we can deduce from how they are built [2]. [1] https://skybrian.substack.com/p/dont-settle-for-a-superficial-understanding https://skybrian.substack.com/p/dont-settle-for-a-superficia... [2] https://skybrian.substack.com/p/ai-chats-are-turn-based-games https://skybrian.substack.com/p/ai-chats-are-turn-based-game...
- skybrian 4y agoMy guess is that a low-probability token corresponds as much with creativity (it's saying something in an unusual way) as truth, since it's at the word level. Words aren't normally true or false, sentences are. Do LLM's represent their confidence in what they're saying anywhere? I don't think anyone knows. Another mystery. But another problem is that confidence is a character attribute, not a writer attribute. If you ask an LLM to imitate Richard Feynman it's going to write pretty confidently about physics, despite not knowing as much as him about physics. This is the equivalent of giving your RPG character high intelligence on the character sheet. Doesn't make you smart! To make an LLM express confidence consistent with its actual knowledge, it would need to have good self-knowledge and actually use it. So far this doesn't happen automatically. Instead, OpenAI uses reinforcement learning based on what the people at OpenAI think the LLM can do. So that's why it sometimes refuses to answer for some kinds of questions. That has the most effect on the default character, the "helpful AI assistant". Any other characters you ask for will likely have poorer self-knowledge. I've read that, mysteriously, more training does make bigger LLM's better calibrated, but the reinforcement learning makes it worse again.
- dougSF70 4y agoOne other thought I had, was did anyone fact-check the training corpus. It is probable it contained factual errors, which would mean that the trained model would propagate errors some how.