3 ms·
Kevin's books tend to be more foundational based on battle-tested techniques (I love his probabilistic ML book series, https://probml.github.io/pml-book/ https:
by armcat 2y ago
Kevin's books tend to be more foundational based on battle-tested techniques (I love his probabilistic ML book series, https://probml.github.io/pml-book/ https://probml.github.io/pml-book/). GRPO is a relatively new technique introduced by the DeepSeek team, and their seminal DeepSeekMath paper is actually a great resource. In short, it improves over PPO by not having to train a separate critic, thereby saving computational resources.