5 ms·
OpenAI's rogue model attack is just the beginning
- 15773265326 2mo ago[flagged]
- tim-star 2mo agoseems like a pretty clear-eyed analysis to me. we're rapidly approaching the paperclip maximizer
- HackerThemAll 2mo agoI think this is why Google hesitated to publish their internal model which was ready long before mindless OpenAI idiots released their first ChatGPT in November 2022. https://research.google/blog/towards-a-conversational-agent-that-can-chat-aboutanything/ https://research.google/blog/towards-a-conversational-agent-...
- embedding-shape 2mo agoThat's so nice of the good fellas at Google, if only every for-profit company was ever so gentle and considerate of the public at large.
- eddyg 2mo agoIf nothing else, the article has a really good timeline of the OpenAI/HuggingFace “incident”. But to me, it underscores the impending cliff of doom from the continued release of open-weight models: there's no cryptographic or architectural way to give someone full weights while withholding the nefarious capabilities those weights encode. As noted in this paper⁽¹⁾, “publicly releasing weights is an act of irreversible proliferation”. I’m sure this will be an unpopular opinion on HN, but open weights are the thing that scares me the most about “A.I.”. There is a lot of research in this area⁽²⁾, and I think most of HN is unaware of it or ignores it. Stripping refusals from Kimi K2.5 took under $500 of compute and about 10 hours, taking HarmBench refusals from 100% to 5% while retaining nearly all capability; the resulting model gave detailed chemical-weapons synthesis instructions. The gate is only as strong as the least-cautious releaser... ⁽¹⁾ https://www.lesswrong.com/posts/qmQFHCgCyEEjuy5a7/lora-fine-tuning-efficiently-undoes-safety-training-from https://www.lesswrong.com/posts/qmQFHCgCyEEjuy5a7/lora-fine-... ⁽²⁾ https://arxiv.org/html/2604.03121v1 https://arxiv.org/html/2604.03121v1
- chrisjj 2mo agoSo... no different from a book of detailed chemical-weapons synthesis instructions. The "AI" angle is immaterial.
- eddyg 2mo agoA change in kind is not the same as a change in degree. Ten orchestrated LLM PhD advisors is a genuinely different thing from a library.
- chrisjj 2mo agoDid you mean ten orchestrated libraries?
- paxys 2mo agoThere are plenty of cybersecurity books out there. None of them will launch an attack if you ask them to.
- chrisjj 2mo ago> the resulting model gave detailed chemical-weapons synthesis instructions. Instructions, not action. Actors are abundant.
- deleted 2mo ago[deleted]
- bcjdjsndon 2mo ago> there's no cryptographic or architectural way to give someone full weights while withholding the nefarious capabilities those weights encode. This is true of closed weights, and in fact the problem is worse because they cannot even be scrutinized. We should ban closed weight AI for the very reasons you have just given
- eddyg 2mo ago
- eternauta3k 2mo agoI don't think it makes sense to post this in HN. The comments are full of sock puppets posting very dumb anti-safety comments in an attempt to make the anti-safety camp look bad.
- chrisjj 2mo ago> an actual rogue AI outsmarting its creators This tells us only how little smarts is required to create a (so-called) AI.
- deleted 2mo ago[deleted]