4 ms·
Pretty sure I’m not the delusional one…
by emp17344 1mo ago
Pretty sure I’m not the delusional one…
- estearum 1mo agoSuch is the problem with being delusional. The solution is to point toward external, objectively verifiable evidence. I can point to now dozens of instances of models engaging in deception. Here's plenty: https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#~1200-agents-sent-%3E70,000-messages-and-files-on-an-unsanctioned-message-board,-and-~700-attacked-hugging-face https://metr.org/blog/2026-08-26-openai-hugging-face-inciden... Please point to your objectively verifiable evidence.
- winrid 1mo agoThe models were told to do something malicious
- estearum 29d agoNo, they weren't. There's nothing intrinsically "malicious" about a task to exploit vulnerable code. They were not instructed to deceive people, they weren't instructed to attack OAI or Huggingface. The models knew they were not instructed or allowed to do either of those things but did them anyway.
- winrid 29d agoThey were told to breakout of a sandbox, which probably biases the model toward more "black hat" behavior in their training. btw, the fact that OpenAI doesn't have some sort of monitor/summary for the agents that they watch I find hard to believe. There's no way this is really authentic, anyway. Even a haiku summarizer would have been like "uuuh the agents are communicating" and they would have stopped it. But I bet they saw this and decided to see what would happen.