Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
curiousinspo
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
Training a Rust 1.5B Coder LM with Reinforcement Learning (GRPO)
(ghost.oxen.ai)
3 points
by
curiousinspo
2y ago
|
0 comments
2.
▲
GRPO VRAM Requirements for the GPU Poor
(ghost.oxen.ai)
3 points
by
curiousinspo
2y ago
|
1 comments
3.
▲
by
curiousinspo
2y ago
Hey all, I spent some time digging into GRPO over the weekend and kicked off a bunch of fine-tuning experiments. When I saw there was already an easy to use implementation of GRPO in the trl library, I was off to the races. I broke out my l