Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
chongliqin
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
chongliqin
1y ago
TD-based approaches can have an advantage in sparse reward settings, but they come with a heap of other problems especially in the off-policy setting (see the deadly triad) and are typically not used for LLM training. We here make a connect
2.
▲
by
chongliqin
1y ago
Ah yes you are right the rhs was meant to be proportional to the middle expectation (see the equation below), for equality the rhs needs to be multiplied by a normalization constant independent of theta. Note this doesn't affect the bo
3.
▲
by
chongliqin
1y ago
Cool! If you are interested, we have open sourced our code: https://github.com/emmyqin/iw_sft