Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
AMavorParker
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
Propel: Breaking the Solver Bottleneck in Task-Generator RL
(vmax.ai)
3 points
by
AMavorParker
4mo ago
|
0 comments
2.
▲
Unix-CTF: Procedural Environments for Unix-Competence Reinforcement Learning
(twitter.com)
2 points
by
AMavorParker
4mo ago
|
0 comments
3.
▲
by
AMavorParker
5mo ago
The teachers never attempt to solve their own problems, only the students solve problems. Regarding the TrueSkill of the teachers, the self-play settings we operate in in this paper are zero-sum competitive which means that the population s
4.
▲
by
AMavorParker
5mo ago
Thanks for your interest! Not necessarily. While the held-out downstream evals showed that 1T-1S setups outperformed larger populations like 4T-4S or 8T-8S on some specific benchmarks, that does not invalidate the motivation for population-
5.
▲
PopuLoRA: Co-Evolving LLM Populations for Reasoning Self- Play
(vmax.ai)
51 points
by
AMavorParker
5mo ago
|
6 comments
6.
▲
by
AMavorParker
5mo ago
We introduce PopuLoRA, a population-based asymmetric self-play framework for reinforcement learning with verifiable rewards (RLVR) post-training of LLMs. Teachers and students are specialised LoRA adapters on a shared frozen base: teachers