Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jacek-123
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
jacek-123
5mo ago
Feels like a training-data artifact. SFT and preference data are full of "here's a cleaner version of your file", not "here's the minimum 3-line diff". The model learned bigger, more polished outputs win. Promp
2.
▲
by
jacek-123
5mo ago
Did you try GradNorm or PCGrad, or was manual task weighting good enough? Also curious about the required-vs-preferred head failing. Was that encoder gradient interference from the other tasks, or a capacity issue in the linear head?
3.
▲
by
jacek-123
6mo ago
We ran this benchmark because we kept seeing the same failure mode: teams fine-tune small models on production traces expecting them to learn their agent's behavior, but the downstream metrics are poor. We tested 5 corruption scenarios