2 ms·
Details on how DS used GRPO for RL rewards https://medium.com/@sahin.samia/the-math-behind-deepseek-a-deep-dive-into-group-relative-policy-optimization-grpo-8a
by mtkd 2y ago
Details on how DS used GRPO for RL rewards
https://medium.com/@sahin.samia/the-math-behind-deepseek-a-deep-dive-into-group-relative-policy-optimization-grpo-8a75007491ba https://medium.com/@sahin.samia/the-math-behind-deepseek-a-d...
- quantumspandex 2y agoThanks!