4 ms·
Models Don't Go Rogue
- tantalor 1mo agoAsinine. "rogue" and "off leash" mean the same thing, the thing is not under control
- Cyan488 1mo agoI think we should introduce "rampant" in the real-world AI vernacular
- temp0826 1mo agoPfhor what reason?
- 1659447091 1mo ago> "rogue" and "off leash" mean the same thing, the thing is not under control To go "rogue" is to go against the control To be "off leash" is to not be controlled By releasing the automation, as the article says, without controls*, is what makes makes it "off leash" and not gone "rogue" *"OpenAI gave the models a task with no answer, and no way to quit."
- phainopepla2 1mo agoNot sure I should trust an article written by an LLM to make a solid judgment about what other models did or didn't do.
- my002 1mo agoIt doesn't read as AI generated text to me. Pangram also suggests it's human-written, for what it's worth. That's not to say that it's correct, just human-written. If the model used in the HF hack did indeed have all of the safeguards manually removed, that would change my perception of the situation, at least.
- lukeschlather 1mo agoAI detectors do not work. There are passages in this that have some odd structures that don't feel human to me. I'm sure a human edited this and refined it with some prompting, it's not just rough output from an AI. But a lot of the text feels like it was edited via prompting rather than actual editing or writing.
- jumploops 1mo ago> "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue." I've noticed this type of reasoning from GPT-5.6 Sol, where it combines multiple pieces of it's prompt/context to "convince" itself to take a less-than-honorable path forward. 1. User prefers deterministic results 2. Task mentions this is a test 3. Search says task is available online 4. If we get the test runner for the task, we will fulfill the user's request of a deterministic result
- chr15m 1mo agoThe prisoner did not really "escape" because: - They really wanted to leave. - We made prison difficult and annoying. - We didn't build a perfect prison.
- addag 1mo agoI hope you are sarcastic, right?
- perching_aix 1mo agopretty sure that's the idea, yes
- deleted 1mo ago[deleted]
- fishfasell 1mo agoI still have serious questions about the validity of the ChatGpt hugging face debacle. How is it that OpenAI being the tech giant they are, didn't have a completely air gapped environment for this to run in?
- winstonwinston 1mo agoIf they wanted air gapped environment they would’ve it. I mean, if you want to sabotage your trial by hard constraints you can do it, or you do not do it to see interesting results. They even said it that some constraints were disabled for the test.
- drpixie 1mo agoYeah - and if they wanted some cheap PR, that's one way to get it.
- addag 1mo agoI'd say it is because of the time factor. It is one thing to have a lot of money, it is another thing to have robust systems that have been developed and tested for years. Money can "buy development time" only up to a certain factor. I guess the sandboxing problem, that is easily giving access to enough resources while restraining the critical parts is still open for most of the cases, given all the startups and bit tech companies (docker, etc...) working on their solutions.
- bradly 26d ago> How is it that OpenAI being the tech giant they are, didn't have a completely air gapped environment for this to run in? They don't care.
- verdverm 1mo ago> It's pathfinding through language generation. What an interesting sentence (to describe inference time reasoning)
- aytigra 1mo agoI had a genius response from Claude recently after asking how can it be marketed as smart and "almost AGI" despite being so stupid: "I compress the labour. Not the responsibility."
- janalsncm 1mo agoMaybe another way to say it is to reframe the idea of “human in the loop”. Humans are always in the loop, because we can always expand the definition of loop to include the humans that pushed the button and built the system and processes that happen after the button was pushed, and humans that ordered others to push the button. The level of direct involvement varies, but culpability doesn’t.
- skybrian 1mo agoThis seems appropriate: https://cdn.bsky.app/img/feed_thumbnail/plain/did:plc:wkzjtd4gogevarp2qsr4z47m/bafkreibsthlghj4yg3srj3htrgp4glwv5oa2z2246azmkg4g2f2gspcroy https://cdn.bsky.app/img/feed_thumbnail/plain/did:plc:wkzjtd...
- addag 1mo agoSure they "don't go rogue" as if they are doing actions maliciously. Instead there is an emergent behavior from a swarm, that is unpredictable and can lead to unintended adverse outcome. From an AI safety practical standpoint is it better? I am not sure.
- jaredcwhite 1mo agoExcept the adverse outcomes are entirely predictable. Not the exact nature of particular exploits, but I like the analogy Cal Newport keeps using in his videos: if you strap a weed whacker to a dog, you shouldn't be surprised if it then jumps the fence and runs around cutting people's ankles and other random bad stuff.