4 ms·
They say it doesn’t need preference data, but it seems to me that this does use preference data - the preferred response is from GPt-4, and the non-preferred re
by tempusalaria 3y ago
They say it doesn’t need preference data, but it seems to me that this does use preference data - the preferred response is from GPt-4, and the non-preferred response is from their model. It doesn’t fundamentally obviate the need to collect a high quality dataset from somewhere else.
In AlphaGo self play, the only external data was grandmaster Go moves that were used in a first pretraining phase of the policy network, and in AlphaGo Zero there was no external data at all. That’s what I would understand as self play really.
Seems to be more efficient than DPO - will try it out to compare
- feoren 3y ago> It doesn’t fundamentally obviate the need to collect a high quality dataset from somewhere else. You'd think this would be obvious: there's only so much worth you can "mine" out of any given data set. We've probably got lots of room left to further refine existing datasets, but clearly you need a high-quality dataset to start with to learn anything at all, right? And yet I wonder ... > In AlphaGo Zero there was no external data at all. That’s what I would understand as self play really. The interesting thing here is the complexity of the strategy -- the crazy amount there is to learn -- vs. the size of the ruleset. The raw materials being mined for strategy here is simply the simple rules, and yet they provide quite an impressive amount of possible learning. Similarly, there seems to be an endless amount that number theorists can learn from the simple rules of integer addition and multiplication. I can't help but think about Chomsky, about the possibility of an implicit grammar, or psuedo-grammar, or grammar-factory, that all humans seem born with. Is there some threshold of examples of human language beyond which an LLM can inductively reason out the "rules" of our innate grammar-factory, and from there, self-play until it has mastered language as well as humans have? As well as humans ever could?