3 ms·
Why not synthesize the two approaches? Reinforcement learning from factual accuracy. Use a language model to run queries against another language model and che
by FuckButtons 3y ago
Why not synthesize the two approaches? Reinforcement learning from factual accuracy.
Use a language model to run queries against another language model and check if it’s hallucinating.
Say we have two language models A and B, A is the verifier and B is being trained.
We give A accesses to a ground truth database, and then we get it to generate questions it knows the answer to based off of its knowledge base.
A asks B those questions and then it verifies Bs output against its knowledge base and we use the veracity of Bs output as the reward function.
- nullc 3y agoCause if you got the facts you train on them, the model memorizes those-- and hallucinates on the facts you didn't teach it in training. :) Call this model A. So you say okay, I'll leave out some facts from training to use to teach it to say I don't know. Now you have model B that is similar to A but on some things that A answers correctly on, B answers I don't know. ... and on some new facts both A and B hallucinate, so B is strictly worse than A-- they both hallucinate but A knows more. Using known unknown facts to train for saying "I don't know" is only useful if it produces a general ability to say "I don't know" against unknown unknowns. And I don't know if anyone has managed to demonstrate that result. It's difficult in general to know what the model does and doesn't know, and LLM isn't a trivia bot-- and it knows tons of stuff that exists nowhere explicitly in the training data (which is why it's useful over and above a verbatim internet search!). It's fun to play with the boundary of LLM knowledge by conversing with it in ROT13 or asking it write with bizarre constraints and watch its intelligence fall away.
- LawTalkingGuy 3y agoBecause the model doesn't contain facts at any point - only words. The size of the LLMs is an easy way to currently demonstrate this - even if you took only the factual statements from all its training material and compressed them, they'd be larger than the model. The model isn't magic thus it can't contain all those facts. The only way it can write a fact is if those words are simply the most likely completions and happen to be right. This means you'd be essentially training it randomly by selecting factual answers. You wouldn't be reinforcing that it gave you a correct fact, just whatever the sentence structure was that you judged to be factual. I think what would happen is that it would start to write very careful statements which would be more likely to be technically correct merely by not being wrong. For example, if you trained by asking questions like what year president George Washington was born it would quickly learn to stop guessing a year because that's got a low probability of being right and those statements would get trained out. It'd probably write something like "Before 1760" because that statement has a much higher likelihood of being right even if it's a less useful answer.