4 ms·
Question about the rule-based rewards (correctness and format) mentioned in the paper: Does the raw base model just expected “stumble upon“ a correct answer /co
by Imanari 2y ago
Question about the rule-based rewards (correctness and format) mentioned in the paper: Does the raw base model just expected “stumble upon“ a correct answer /correct format to get a reward and start the learning process? Are there any more details about the reward modelling?
- leobg 2y agoGood question. When BF Skinner used to train his pigeons, he’d initially reinforce any tiny movement that at least went in the right direction. For the exact reasons you mentioned. For example, instead of waiting for the pigeon to peck the lever directly (which it might not do for many hours), he’d give reinforcement if the pigeon so much as turned its head towards the lever. Over time, he’d raise the bar. Until, eventually, only clear lever pecks would receive reinforcement. I don’t know if they’re doing something like that here. But it would be smart.
- fspeech 2y agoSince intermediate steps of reasoning are hard to verify they only award final results. Yet that produces enough signal to produce more productive reasoning over time. In a way when pigeons are virtual one can afford to have a lot more of them.
- whimsicalism 2y agothey’re not doing anything like that and you are actually describing the failed research direction a lot of the frontier labs (esp Google) were doing
- whimsicalism 2y agoyes, stumble on a correct answer and also pushing down incorrect answer probability in the meantime. their base model is pretty good
- stri8ted 2y agoIt seems a strong base model is what enabled this. The models needs to be smart enough to get it right at least some times.
- pama 2y agoThe prompt in table 1 makes it very likely that the model will use the correct format. The pretrained model is pretty good so it only needs to stumble upon a correct answer every once in a while to start making progress. Some additional details in the Shao et al, 2024 paper.
- nialv7 2y agoYes and no. In their paper they said they trained two models. One is purely RL based (R1Zero). So this one is trained like you described, i.e. it has to stumble upon the correct answer. They found it to be good but has problems like repetition and language mixing. The main R1 model was first finetuned with synthetic CoT data before going through RL IIUC.