Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ziaowang
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
ziaowang
2y ago
Texts in the wild used during pre-training contain lots of biases, such as racial and sexual biases, which are picked-up by the model. During RLHF, the human evaluators are aware of such biases and are instructed to down-vote the model resp
2.
▲
by
ziaowang
2y ago
This understanding is incomplete in my opinion. LLMs are more than emulating observed behavior. In the pre-training phase tasks like masked language model indeed train the model to mimic what they read (which of course contains lots of bias
3.
▲
by
ziaowang
2y ago
Can you provide a link to the comment? R1's technical report ( https://github.com/deepseek-ai/DeepSeek-R1/blob/main/DeepSee... ) says the prompt used for training is "<think> reasoning proc
4.
▲
by
ziaowang
2y ago
Agreed. If it's useful, why not scale up the electricity and water supply, and make the latter sustainable.
5.
▲
by
ziaowang
2y ago
Though FAANG offers are usually more attractive than startups (considering pay level and stability), some startups could be more selective since they couldn't afford to hire the wrong candidate.