3 ms·
So this is a good idea because it can create vastly more training data for a model to learn from. However, it seems likely that these models are going to halluc
by williamtrask 3y ago
So this is a good idea because it can create vastly more training data for a model to learn from. However, it seems likely that these models are going to hallucinate like crazy. As featured in the documentary, the AlphaGo program struggled with hallucination, and it performed self-play in a tiny world based on a perfectly rigid and exactly correct set of rules. LLMs already have tons of false corners and edges in their logic about the world, and this seems like it has the potential to spread those all around.
Hard to say — these things can be difficult to predict. I can see this working but there'll probably be some ratio of training data - self-play that we have a hard time getting past because it's a difficult-to-control form of extrapolation.
- sdenton4 3y agoIt's a very promising idea, though. Judging content is much easier than creating content, after all.
- williamtrask 3y agoIt is a promising idea, but it falls prey to the same tautology as in the game-playing agent days. In order for simulation to be useful, you need a really robust model of the world, but if you had a really robust model of the world, you wouldn't need the simulator. Simulation is really really good for one thing: learning a policy that can find a particular corner of the search space as fast as possible (such as a wining state in a Go game). But simulation is not good at actually generating the space. It's going to extrapolate mistakes really badly. Also, while humans are better at judging (look at that fake thing!) than generating (drawing a realistic photo), I think you may find that — despite our abilities — detection is actually quite a bit harder. As an example: "John Doe is dead." This was very easy for me to create, but it's quite difficult for you to judge whether or not it is true due to a variety of factors (which John Doe, am I being honest, when was the last time you saw John Doe, do you know anyone who knows John doe and could check, perhaps John Doe had a twin who died, etc.)
- sdenton4 3y agoYour problem is harder than it needs to be; we can convert it to a closed problem, by requiring citations. Then the question is whether the statement is properly supported by the citations, which is far easier to evaluate.
- williamtrask 3y agoInteresting direction! Let's assume the citation is me. Does that make content generation harder than verification (judgement)?