Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zimablue22
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
zimablue22
2mo ago
Yeah. This work is over claiming the novelty quite a bit.
2.
▲
by
zimablue22
2mo ago
Any of the Best of N papers that exploded in popularity after GRPO. Likelihood is not fundamental to the spirit of GRPO, any exploratory mechanism would work. That sequential LLMs have a step-wise probability is convenient but not critical
3.
▲
by
zimablue22
2mo ago
You're right about the original GRPO proposal, but there are simplified variants that do just use best of K sampling. GRPO (or GRPO like approaches) for diffusion/flow matching similarly can be likelihood free.
4.
▲
by
zimablue22
2mo ago
That's fair. On the other hand, Minibatch OT (optimal transport) was one of the more fundamental advancements early on in flow matching & rectified flow models. This best of K approach effectively discards matches that would otherw
5.
▲
by
zimablue22
2mo ago
This is just GRPO (proposed by DeepSeek), which similarly samples many plausible generations, selects the best of K, and trains that sample. Minibatch OT in flow matching also has a very similar mechanism, where samples from a noise distrib