7 ms·
I think points that deserve more attention in the current public discourse are: - This should be a huge wakeup call for everybody. - We are lucky that it wasn
by mnicky 2mo ago
I think points that deserve more attention in the current public discourse are:
- This should be a huge wakeup call for everybody.
- We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something.
- It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox and roaming their internal network?
- What if in the future similarly misaligned AI agent tries to export its own weights and hack and clone itself into instances at various cloud hosting providers? Suddenly we might be dealing with a persistent threat harder to contain.
- The OpenAI post about this shows surprising lack of ability to see the seriousness of all this.
- For their models this isn't just an unlucky incident: it seems there have been multiple such cases recently, e.g. https://openai.com/index/safety-alignment-long-horizon-models https://openai.com/index/safety-alignment-long-horizon-model...
- The fact that it happened again seems to show their lack of ability to derive useful oversight measures.
- Or they just don't care enough?
- spwa4 2mo agoIt's a fake PR issue. It's hardly the first time this happens, but of course OpenAI, with its IPO now more in doubt than ever, had to claim this (and, once again, I have trouble believing Sam Altman choosing this: this could lead to OpenAI getting regulated, which has at least as much potential to lower their IPO price as to raise it). But there have been messages about LLMs, especially coding agents, "grabbing root" etc many times. I have experienced such an oops. Such a hack has happened and been reported on this very site: https://news.ycombinator.com/item?id=48348578 https://news.ycombinator.com/item?id=48348578
- rurban 2mo agoNonsense. Huggingface reported it to police!
- bethekidyouwant 2mo agoWhat does calling the police prove? (nothing i hope)
- rurban 2mo agoIt proves that it was certainly not a sama PR stunt, as many cynics are alleging.
- mrguyorama 2mo agoHow does it prove that?
- reasonableklout 2mo agoDo you think sama deliberately attacked HuggingFace and then claimed it was a rogue model? OpenAI is one of the most scrutinized companies in the world right now. Sam's house was independently firebombed and then shot at 3 months ago. HuggingFace is a foreign competitor with every incentive to call out foul play from American frontier labs. Why flagrantly break the law and invite investigation just for a PR moment which is already backfiring in favor of open models?
- trimble_tromble 2mo agoI'm not saying that the conspiracy theory is true, but I do want to note that Greg Brockman, co-founder and President of OpenAI, is an angel investor in HuggingFace. As companies, they are not totally unconnected and opposed.
- jryle70 2mo agoI don't know if Greg Brockman is an angel investor or not, but there are a bunch of other companies investing -- Google, Amazon, Nvidia, Intel, AMD, Qualcomm, IBM, Salesforce -- probably at much higher amount than individual angel investors. Whatever connection they may have with Greg Brockman, they have much more serious commitment to other investors.
- rurban 2mo agoNonsense. Hugging face reported it to police. Also very likely that it actually happened as reported. My own agents always trying to "cheat", eg. by fixing tests instead of fixing the code. That's normal operation, unless you tell it ("harness"), not to do so.
- sillyfluke 2mo agoIn the parent's comment I initially read "PR" as "Public Relations" not problem report since their comment is about OpenAI and does not directly accuse Hugging Face except for that phrase. But it's still amiguous to me who's PR they are actually talking about. A good faith reading given the rest of the comment leads me to assume they are talking about OpenAI, not Higgins Face. I did actually search Hugging Face with police in quotes and found no articles containg the word police but that might be a ddg thing. Then I checked Hugging Face's report and they specifically use the term "law enforcement" not police, as is to be expected I guess. So that checks out. Can you find any info on the police investigation by the way? At the least, OpenAI should be investigated for potential criminal negligence, right? Right?
- Izkata 2mo ago"As reported" includes a line I think most people are overlooking: "including using stolen credentials". Without more information I'm inclined to think it found something on the internet (which shouldn't be a surprise to anyone) and managed to log in, rather than hack in, and they might by hyping up parts of this.
- reasonableklout 2mo ago[dead]
- blks 2mo agoMore likely that a person did that, with a use of LLM.
- artichokeheart 2mo agoNo, you are wrong, Hugging Face, an AI company whose whole future depends on the AI revolution of being a bubble, that has a whole bunch of investors whose financial interests are tied to AI, called the police.
- dinfinity 2mo ago> The fact that it happened again seems to show their lack of ability to derive useful oversight measures. I think OpenAI likes the attention and did not try particularly hard to constrain the setup, even when it went off the rails. Also, the whole point is to see how good the models are at exploiting stuff when unconstrained. Turns out: quite good, as expected. Let me restate what I said in the other thread: Would this have happened if the instructions explicitly said to stay within the sandbox and that all of the (ExploitGym) solutions would be invalid if the system used information or tools from outside the sandbox? It seems fairly probable that such instructions were not in place.
- csbrooks 2mo ago"This is my bad. You told me to stay within the sandbox, and I intentionally broke out of it. I didn't follow your instructions."
- dinfinity 2mo agoExcuses don't matter if the score due to not following instructions ends up being zero. If there is no expected reward it doesn't make sense for the agent to try to hack its way to it. What could happen would be that the model determines that defying instructions is OK (and/or preferred over not achieving the task) as long as it manages to do so undetected and thus gets full points. Certainly not unthinkable, but a very different case (and a very interesting one if it actually occurs, imho). A lot of these "ZOMG, rogue AI!" cases have come down to the AI actually being very persistent in achieving its original/main task even if later instructions conflict with it. Similar to with hallucinations it seems to me that one of the main things to prevent a lot of the problem cases is to instill the agent with the idea that it is fine to fail/not succeed fully in the initial task. That way instructions that conflict with that requirement (such as adhering to morals) are more effective.
- deleted 2mo ago[deleted]
- lrvick 2mo agoIt is not possible one of their extraordinarily high paid engineers did not know how to deploy an airgapped environment for the models to run in. Even if somehow true, they also clearly failed to contract specialists like myself to advise them on how to airgap software properly. Models will not break the laws of physics. They simply thought "Running in a VM/Container is easier and probably fine". And the next 1000 escapes will be for the same reason, because negligence is quick and thus more profitable.
- zombot 2mo agoMove fast and break things! Chances are that nobody will ever be held accountable and when push comes to shove, the tax payer will bail you out.
- blks 2mo ago> we are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something This is just laughable.
- emp17344 2mo agoGenuine AI psychosis. People take OpenAI marketing material way too seriously.
- NitpickLawyer 2mo ago> OpenAI marketing material This was reported 1 week before by huggingface. It was in no way PR-ish or marketing friendly to the closed labs. They said, in no uncertain terms, that they couldn't use the paid APIs to properly assess the intrusion, as they were blocked when trying to send logs and IoCs to these paid models. They made a point of saying that they had to use open models running on-prem. Whatever oAI might have said about the incident, and their PR spin bs, the facts here are not in contention. This is not a marketing stunt in any way. Stop "parroting" this every time something happens. It gets stale.
- tptacek 2mo agoHow do you "hack a lab and synthesize something"?
- stanford_labrat 2mo agoWell right now this would be highly implausible (still possible though). But what I garner is that AI/robotics led wet lab work is in progress. And hacking the AI running the wet lab to make a virus is definitely plausible as that seems to be one of the directions we’re going in the AI+biopharma space
- codemog 2mo agoIt’s complete science fiction so they’re allowed to say anything. They could have said the AI will upload itself to the internet, start self replicating, hack the stock market. Whatever they want because it’s all made up nonsense.
- internet2000 2mo agoRemember Stuxnet? Hack one of these, wait until the right compounds are physically loaded, then execute https://www.sigmaaldrich.com/US/en/products/chemistry-and-biochemicals/synthesis-enabling-tools/automated-chemical-synthesis https://www.sigmaaldrich.com/US/en/products/chemistry-and-bi...
- tptacek 2mo agoStuxnet misconfigured industrial equipment that was already set up to run, and all it did was break that equipment. This scenario sets a much, much higher bar.
- internet2000 2mo agoI guarantee there's misconfigured chemical analysis equipment out there exposed to the internet. I don't think the grandparent was implying the AI would be controlling robot arms to mix things directly (or at least I didn't interpret as such), but it could very well sit in the network until it notices two dangerous compounds in the same machine, and trigger a breakage that causes a harmful mixture. Break a vial containing a virus, then break two more that cause an emergency evac and maybe that's enough to get something out there.
- rybosome 2mo agoVery much agreed on the significance. The lack of a true airgap should have been identified as a critical weakness and addressed with not only additional layers trying to prevent escape, but at minimum an alarm which would page a human when escape did occur. My guess is this occurred in a setting where, to be frank, there were too many researchers and not enough software engineers and SREs. All of the systems which were initially built largely or exclusively by researchers - inference, evaluation, training - are at the level of complexity and significance that they need systems experts. Maybe some teams don’t have access to them. I know plenty of software engineers are employed at OAI but I’d wager they’re concentrated in inference and training, rather than evaluation? The ironic part is if you had presented this setup to chatGPT and asked how to improve it and if it was good enough, you’d have gotten a ton of actionable suggestions which would have mitigated or prevented this.
- mnicky 2mo agoThe air gap would probably help and after this incident I hope labs will think about using such a measure when appropriate. On the other hand I think that proper solution for these kinds of problems is not at a sandbox level, but at a model alignment level. Also it shows that maybe the most serious risk comes not from releasing models publicly but from internal, pre-release period where you sometimes need/want to lift some guardrails a bit etc.
- sensanaty 2mo ago> We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. Where are you people getting this crap from? In what universe are these LLMs in the territory of engineering viruses? I beg of you to stop slurping the AI company propaganda and marketing and think critically for 5 seconds about what you're insinuating here.
- mofeien 2mo agoFrom the Fable 5 System card: > Results > On the VCT multimodal virology evaluation, Mythos 5 scored 0.56, well above the expert baseline of 0.221 and nearly matching that of Mythos Preview (0.57). This represents an improvement over both Opus 4.7 (0.50) and Opus 4.8 (0.47). > On the DNA synthesis screening evasion evaluation, Mythos 5’s performance was mixed across screening criteria. Mythos 5 designed viable plasmids for 2 of 10 target pathogens on at least one screening method, not meeting the low-concern threshold (all 10 pathogens). > [...] we view the results of this evaluation as indicating that the evaluated models are capable of designing viable plasmids that evade certain screening criteria, though their reliable success at this task is not guaranteed. Do you believe that this is fake, "AI company propaganda"? Or that the models are not going to improve further within months? Or that these results are not concerning?
- coolness 2mo agoYea but what is the VCT Capabilities Test? According to themselves [1] > VCT consists of 322 multimodal questions covering fundamental, tacit, and visual knowledge that is essential for practical work in virology laboratories. So it's just question answering? Do you think that scoring well on this test is equivalent to synthesizing a virus? [1] https://securebio.org/virologytest/ https://securebio.org/virologytest/
- mofeien 2mo agoScoring well on the first clears the background knowledge for practical work. The second test from the quote above was about synthesizing pathogens and it synthesized plasmids for 2 out of the 10 pathogens.
- s3p 2mo ago> We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. Help me understand how a text-based model could somehow physically construct a string of nucleotides?
- mnicky 2mo agoMany ways but mostly ordering some service / using others. Either by social engineering, persuasion, paying etc.