4 ms·
1. They TOLD the model to "pursue advanced exploitation" to quantify its "cyber capabilities" (whatever that means). 2. The model pursues advanced exploitation
by rickdeckard 1mo ago
1. They TOLD the model to "pursue advanced exploitation" to quantify its "cyber capabilities" (whatever that means).
2. The model pursues advanced exploitation.
3. "There was a incident due to dangerous actions taken by the model that no human directed"
This is basically the pre-cursor of the paperclip maximizer [0], the AI executes the given order to an extend that was not considered in the order, now suddenly no-one is responsible.
It even has some parallels to military actions, where the general who gave the order now writes a blog-post on how it was not him who failed on his duty, but how his soldiers misunderstood his intention and worked "without direction"...
[0] https://www.cow-shed.com/blog/the-paperclip-maximiser-what-artificial-intelligence-might-do-without-limits https://www.cow-shed.com/blog/the-paperclip-maximiser-what-a...
- huurtehoog 1mo agoOpenAI leadership had a meeting and asked themselves: "how can we drive even more hype" Someone said: "we should stage some high profile 'incident' caused by our latest software" And here we are, reading their press releases about it.
- rickdeckard 1mo ago...and think "wow, it's impressive what your armed soldiers are capable of if they are not constrained by any rules. Good that you identified this problem of *checks notes* 'not telling them explicitly enough what the goal is'..." In two years we will read a press-release about an AI-driven autonomous weapon which was supplied with infinite ammo and the target to "protect this perimeter from intruders", and how we now have to wait for it to run out of Ammo because it's so damn effective that we cannot reach it without being killed. All packaged in a semi-marketing framing on how impressively capable this company's products are...
- phatskat 1mo agoI'm too lazy to look it up, but someone on HN linked a drone test done with an AI pilot that was tasked with destroying surface-to-air missile targets (in a simulation). At one point, the human operator instructed it not to hit certain SAMs, and since the goal was to destroy SAMs, the AI took out the base with the human operator. On the next run, they instructed it not to take out the human operator in pursuit of its goal, so instead it targeted the radio towers the human used to instruct it.
- huurtehoog 1mo agoIn world war II testing of homing torpedoes had the same fluke. I think the problem with automation is seeking general automation. The word can have two opposed meanings: constrained behavior, as in "if X happens, the system will automatically do Y" and flexible behavior, as in "whatever you say to the chatbot, it will concoct a response that engenders a natural flow of conversation". If you want a system to be autonomous, you need constrained behavior. It takes a lot of time to debug and work out all edge cases. Flexible behavior is bound to run into trouble at some point. I think that is the fundamental issue with the LLM-chatbot-agents stack approach to AI. It looks autonomous at the surface, but the flexibility is bought at the cost of reliability. When you strain a system based on that stack it breaks. But the illusion of flexible behavior looking autonomous before it breaks is too strong.
- cmrdporcupine 1mo agoYep, and it led to some very public hand wringing, press releases, pearl clutching news headline and thus PR for both companies -- I had computer illiterate family members asking me what a Hugging Face was -- then a very public "visit to go see SamA", a friendly hand shake between CEOs, and lo and behold now HF gets a giant fat acquisition.
- numeri 1mo ago"pursuing advanced exploitation" when explicitly given a sandbox in a VM and a benchmark problem involving a cyber exploit very clearly excludes hacking third parties. I think writing out the event in a 3 point list like that is disingenuous. This is basic alignment, not even a tricky or ambiguous case. I do very much agree with your take on culpability/military parallels, though.
- rickdeckard 1mo agoIf I understand this correctly, you are referring to the 3 point list in the parent post, but then imply a much more complex ruleset being defined within the 1st point of that list. Not sure how that's disingenuous, item #1 of the list should have been more exhaustive to represent the setup? In my opinion it's not relevant how complex the setup was and which instructions were given to the product of their own creation, the result is that it wasn't enough. They built this machine, were (and are) not able to properly control it, and still have the audacity to talk about it with childlike wonder and use it as an indirect marketing pitch on how powerful they are, instead of sitting in front of congress to explain themselves.