Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
tamassimond
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
How to Train an LLM to Do Proofs: Beyond Verifiable Rewards
(tobysimonds.com)
3 points
by
tamassimond
1y ago
|
0 comments
2.
▲
The Cost of Winning:How RL Training on Poker Leads to Evil LLMs
(tobysimonds.com)
2 points
by
tamassimond
1y ago
|
1 comments
3.
▲
The Hidden Cost of Winning:How RL Training on Poker Degrades LLM Moral Alignment
(tobysimonds.com)
8 points
by
tamassimond
1y ago
|
0 comments
4.
▲
by
tamassimond
1y ago
To clarify we use an Elo ranking system to update models scores, so if you loose to a higher rated story you don't loose as much Elo ranking. Definitely agree with LLM judge criticism though it's still an open questions of how we
5.
▲
by
tamassimond
1y ago
Yeah as we mention in the blog it's really hard to eval on short passages. If you go on the Github can see longer stories where the change is more noticeable. Both those stories are from the same prompt
6.
▲
by
tamassimond
1y ago
Thanks for feedback added MIT license
7.
▲
by
tamassimond
1y ago
Just to clarify slight misunderstanding the variants without parent ID aren't from the initial batch it just didn't carry over to the next batch. You can see "897ccd25-4776-4077-a9e6-0da34abb32a4" emerges from batch 5.
8.
▲
AlphaWrite: AI that improves at writing by evolving its own stories
(tobysimonds.com)
80 points
by
tamassimond
1y ago
|
159 comments
9.
▲
Self Rewarding Self Improving: Autonomous LLM Improvement
(arxiv.org)
28 points
by
tamassimond
1y ago
|
0 comments
10.
▲
LLMs for Engineering: Teaching Models to Design High Powered Rockets
(arxiv.org)
124 points
by
tamassimond
1y ago
|
45 comments
11.
▲
Text to RL: Extracting High-Quality RL Questions from Text
(tufalabs.ai)
1 points
by
tamassimond
2y ago
|
0 comments