4 ms·
I _think_ that's a factor of using rule based reward functions and not actually a feature of GRPO? The original formulation of GRPO from deepseek math uses a ne
by maxrmk 2y ago
I _think_ that's a factor of using rule based reward functions and not actually a feature of GRPO? The original formulation of GRPO from deepseek math uses a neural reward model that is trying to predict human rankings of responses, and in that configuration won't see 0 gradient updates.
Flipping it around, if you swapped out the neural reward model in PPO with a reward function that can return zero, I thiiinnkkk it would be able to produce zero (or very low) gradient updates.
I'll be the first to admit that I don't know enough about the space to say though. I'm still a beginner here.