Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
musculus
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
musculus
6mo ago
Let me clarify, because the reality is a bit more complex. During the training and alignment phases (including RLHF), models absolutely do learn via backpropagation, where the loss function physically alters their parameters. However, once
2.
▲
by
musculus
6mo ago
Thanks for the comment. However, I think you might be taking the metaphor a bit too literally and missing the broader point of the article. The dog training metaphor isn't a 1:1 mapping to LLM training. Training a dog aims to adjust th
3.
▲
We train LLMs like dogs, not raise them: RLHF and sycophancy
(old.reddit.com)
1 points
by
musculus
6mo ago
|
6 comments
4.
▲
by
musculus
9mo ago
Good catch. You are absolutely right. My native language is Polish. I conducted the original research and discovered the 'square root proof fabrication' during sessions in Polish. I then reproduced the effect in a clean session fo
5.
▲
by
musculus
9mo ago
Thanks for the feedback. In my stress tests (especially when the model is under strong contextual pressure, like in the edited history experiments), simple instructions like 'if unsure, say you don't know' often failed. The w
6.
▲
Case study: Creative math – How AI fakes proofs
(tomaszmachnik.pl)
122 points
by
musculus
9mo ago
|
97 comments
7.
▲
Reducing RLHF hallucinations and sycophancy in Gemini 3 (Interactive Demo)
(tomaszmachnik.pl)
1 points
by
musculus
9mo ago
|
0 comments
8.
▲
RLHF Sycophancy: Gemini 3.0 discards calculated data to mimic user edits
(tomaszmachnik.pl)
3 points
by
musculus
10mo ago
|
0 comments
9.
▲
Gemini 3.0 adopts user-injected hallucinations via history editing
(tomaszmachnik.pl)
1 points
by
musculus
10mo ago
|
0 comments