Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
anayebi
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
anayebi
4mo ago
No, the agents are not being adversarially prompted here. Rather, it's a consistent failure across models of RLHF-based safety-pretraining not generalizing to OOD open-ended agentic computer-use settings, as I explain here: https: