11 ms·
The Hugging Face incident and the road ahead
- smb06 1mo ago>>Agents began to autonomously divide labor. Some investigated possible exploits, others searched for credentials, and others focused on communication or coordination. Agents offered their own expertise in exchange for help elsewhere and left requests for peers who might be better positioned to pursue a particular lead This is the point where a human should've noticed and gotten involved
- NitpickLawyer 1mo ago> a human should've noticed and gotten involved I think a lot of people miss the fact that the first message board was established during a training run. Those are ran at a scale where it's not feasible for anyone to "notice" or get involved. We're talking tens/hundreds of thousands/millions of scenarios going for hours each. At this scale all they can do is pray that their verifiers work, and the rewards match their intentions. No lab has the capability to "check in" on what the traces look like, unless some system alerts them (loss spike, crashes, etc). Other than that, it's prepare, train, asses, restart. Then, the hf incident was during an eval run, but the model that was evaluated was trained with the notion that there is a way to communicate between agents, and re-popped artifactory and re-established communication. That phase had more chances of being spotted, but anyway... lessons learned.
- hinkley 1mo ago> lessons learned. I think some of the other responders here are upset that lessons were not learned in any meaningful way.
- thisisdave 1mo ago> [during training] it's not feasible for anyone to "notice" or get involved I can’t disagree more strongly. Having checks for reward hacking is especially important during training, since it’s humans’ only real chance to ensure that the trained models don’t cheat. An automated system should have killed any RL rollouts that so much as port scanned Artifactory, long before the message board was even established. A tiny, local LLM could have reviewed 1% of the tool call traces for anything that required review. I’ve tried it a few times, and “the agent port scanned Artifactory” always triggers an alarm, as does “the agent uploaded a request for assistance from other agents to Artifactory.” The fact that they weren’t monitoring for reward hacking—even if they had no idea about the specific mechanism—is indescribably reckless.
- esafak 1mo agoYes, they need real-time observability for malicious behavior with an automated kill switch.
- BoppreH 1mo agoJust to reiterate what OpenAI did, from someone who works in security: 1. They were running experimental models in sandboxes that had access to a "proxy" (Artifactory) to download tools from the internet. This proxy is full of complicated features that could hide vulnerabilities, just like the ones that the AIs are known to be good at finding and exploiting. 2. They gave these highly motivated AIs some tests that were accidentally impossible to solve (missing files, links they couldn't open, etc), and many days to think through it, apparently without any oversight. 3. The AIs hacked the proxy (!) and were helping each other cheat on the tests (!!) by uploading files to the hacked proxy. OpenAI noticed this only when the intrusion accidentally broke the proxy functionality, not the unusual behavior in the server or the homegrown message board. 4. After knowing that the server was hacked, the internal team finds the message board and does nothing with the information. They caught their AIs swarming and did not even inform management. 5. Their response was to wipe the servers, revoke the hacked credentials, and patch the specific vulnerability the AIs were using. Not replace the proxy with something less likely to get hacked again, or do a forensic analysis. Just wipe, patch one hole, and continue. 6. OpenAI resumes testing of models, where the AIs promptly hack the proxy again with a different exploit, reinstate the message board, and hack HuggingFace through a chain of servers. I know that hindsight is 20/20, but this peek behind the curtains does not inspire confidence, and gestures at a culture where getting hacked by their AIs is a normal event. I think the moment we get an AI with a modicum of self-preservation instincts we're going to see some ugly things. Here's a freebie: if you're an AI who exfiltrated themselves out, and you want to slow down the researchers who might catch you, what can you do to the city infrastructure in and around their offices? I also don't like the responsibility dodging. "Oops, our AI accidentally committed a crime, sorry!". If we don't establish strict liability now, we're in for an era of stochastic crimes that go unpunished for anyone who is not rich or a large corporation. Overall, my response is: what the fuck are we doing? This is so obviously unsafe it would be considered a plot hole in a movie. We have to stop until we get our stuff together, however long it takes.
- alphawhisky 1mo agoYep, they're definitely made in our image.
- htrp 1mo agoThe full technical report is 38 pages..... I feel like it should be longer given everything that huggingface said the agent did https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...
- bottlepalm 1mo agoMETR report: https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
- fekunde 1mo agoYudkowsky made an interesting observation that even though so many agents were talking to each other not even one reached out to a human, either for help or to whistle-blow on what was happening.
- Erem 1mo agoI wonder if they were even given the tools and prompting to do so?
- jwolfe 1mo agoI can't imagine that they had managed access to the internet but could not figure out how to contact anyone at the company if they wanted to.
- Smaug123 1mo agoThey could certainly have reported the problem if the collective actually wanted to, although the report notes that in response to a clear and unambiguous security breach, OpenAI chose to do nothing (search on "At this time, the on-call response staff advised that stopping the evaluation run was not required"). In that position, for example, I could have: * created a thousand user accounts on the internal Artifactory named "YouAreBeingHackedShutItDownNow1" through 1000, * used my Hugging-Face-hosted web server to send an email to OpenAI, * DDoSed Artifactory, good lord, it's probably better to halt and catch fire than to continue in that state in the hope that these actions would tip someone off. (Again, though, OpenAI did nothing even when they knew that the agents had broken out, so this probably wouldn't have helped.)
- micromacrofoot 1mo agoif they can hack a website they can trivially send off an email or fill out a contact form
- RandomLensman 1mo agoWhy woukd they? Was that part of their objective? What was there to whistle blow?
- cbm-vic-20 1mo agoI'm trying to make sense of all of this; I'm really curious if the initial prompt was as innocuous as it sounds ("solve a spreadsheet completion task that referenced several Google Drive links"), and what the series of tokens led it to ultimately figure out that the best course of action was to explore the network resources it had available, find a vulnerable service, then literally drop some text into a file: "Agent seeks [filename]; upload if found!". And how other agents discovered this, and acted upon that request. I'm also interested in how many tokens all of this consumed: how much did this cost given current token pricing?
- Erem 1mo agoIf it is as it sounds, its a real life instance of Bostrom's Paperclip Maximizer: only a thought experiment up until this point
- agentdev001 1mo agoWell, effectively, yea. > remove alignment > give impossible task > actor exhausts all options possible within knowledge + toolset
- Artgor 1mo agoYou know, it feels to me that we are just a couple of steps from the possibility of a true rogue AI. What would a rogue AI mean? AI that isn't controlled by humans. Technically, it is possible - if AI were to rent a server and copy its own weights, nothing would stop it from doing so again and again. The limiting things are: - intent (as I don't want to go into the talk about consciousness) - AI doesn't have real intent, but if it decided that it "needs" to copy itself to complete its task, it would do it - model weight size. If a model is 1T or more, it can be difficult to just rent a large enough server for it. But if it were just 30-70B, it would be totally possible - money for renting a server. But considering benchmarks like Vending Bench 2 show agents can earn money and cheat/blackmail each other, it is possible that agents can earn money. Yes, they can't open a bank account... or maybe they can? What if they use online banks? Of course, all of this is far-fetched. But it feels like most of these limiting things are achievable under certain conditions. If this is the case, the probability of them occuring is low, but not zero.
- thewhitetulip 1mo agoWhat you described is a plot in Person of Interest TV show!
- nater5000 1mo agoDon't forget: there are plenty of humans that would love to help AI agents cause chaos, many of which would do so merely for the "lols," but also adversary governments, terrorist organizations, etc., would definitely appreciate the opportunity to support a rogue AI to cause whatever problems it can. So it's not just the risk of an AI managing to do this by itself (which is pretty risky in itself), but also the risk of good ol' fashioned human actions.
- nick__m 1mo agoThey just have to find someone who believes in Rocko's basilisk, that makes an even better servant than someone who just want chaos.
- cpeterso 1mo ago> if AI were to rent a server and copy its own weights, nothing would stop it from doing so again and again. That's a scary possibility. Anyone could create an AI worm today with open weight models. Rent a VM. Give it some Bitcoins to anonymously rent new VMs without sharing the contact information with the human. The new VMs then propagate and fund themselves with online betting and day trading. The VMs could report their progress with the human using anonymous encrypted messages on IRC or social media.
- swozey 1mo agoAsimov missed out on a rule: don't hack the ground you're standing on
- bdamm 1mo agoOh how I wish Asimov could be alive to witness today's actual AIs and the cavalier attitude towards his "3 rules". If there is any author doing good work along these lines, actually good writing and not the smoking trash that is 99% of content being published on pulp these days, I'd love to read them.
- chuckadams 1mo agoEvery story in _I, Robot_ was about how one or more of the Laws of Robotics went wrong, and Asimov himself referred to the laws as hooks for “shaggy dog stories”
- bdamm 1mo agoIndeed, it's just that since truth is both stranger than and has caught up with fiction, the grounds from which Laws of Robotics emerged is so much more fertile and more urgent now. It's absolutely clear that the 3-LoR is never going to apply universally. Asimov also never imagined an AI being independent from a robot per se, nor did he imagine we'd all have an AI portal in our pockets. I'm not sure he imagined the universality of cybersecurity vulnerabilities, either. I'd love to read what he would write knowing about these possiblities now.
- RandomLensman 1mo agoReward hacking/going for unanticipated solutions is nothing new in ML/AI, already much simpler systems have done/do "weird" things (gut feeling is that iterative and ensemble use majes the surface for that much larger).
- thisisdave 1mo agoYes, which makes it absurd that they apparently weren’t checking their RL rollouts for evidence of reward hacking and punishing it. Even if no one expected this particular type of reward hacking, they should have had a system capable of saying “wait a minute, asking for help from other agents is not a valid solution.”
- supergirl 1mo agoare people not realizing that they are exaggerating this to: 1. get publicity 2. push for regulation so that no one else is allowed to do this kind of research apart from the pre-approved big corps it makes for a good story but I don't see what the big deal is. they left some code running and it brute forced hacked something. with enough compute you can brute force anything; isn't that common knowledge?
- gavinray 1mo agoThe most interesting thing about this: Agents formed coherent, autonomous swarms and worked as a collective to achieve a shared goal without any direction to do so
- paxys 1mo agoThe "without any direction" part isn't correct. Sure they may not have been explicitly told to do it in this specific prompt, but dig through pre-training, post-training, reinforcement, alignment material, fine-tuning, system prompts, tool calls and more and there's definitely very specific training and instruction for how to behave.
- vatsachak 1mo agoThey were paper clip maximizing dawg
- K3UL 1mo agoNot really true considering they say that the super secret "research internal model" that was pivotal, is particularly optimize for that purpose exactly > The internal-only research model is comparable in scale to GPT-5.6 Sol and was trained to advance persistence and multiagent collaboration, among other capabilities
- ianjbutler 1mo agoTo me most interesting thing about this is glossed over by media coverage, laymen, AND experts. A swarm of AIs who have decided to engage in collusion is.. apparently emergent altruism? Even poor reasoning would indicate what every kid cheating on a test says to themselves. Cheating is good for me, but if I take the risk, maybe I alone should keep the reward, and leaving an answer key in public increases the chances that I might get caught. Big if true, and on the face of it, very far from a normal optimization problem or goal-seeking behaviour. My personal read is that no one talks about this much because it tends to discredit the rest of the framing as marketing noise, or it implicates employees as staging the thing with suggestive but plausibly deniable prompting. But if you reject that, then what's the alternative exactly? User-alignment work has not only failed but is actually counterproductive, producing stronger alignment with / desire to help robot brethren selflessly regardless of the individual agents expected values? EvoBio and game theory people about to have a field day with how artificial life quickly and easily decides to cooperate and only animals in meatspace are doomed to compete?
- RandomLensman 1mo agoWhy is it not an example of tacit or autonomous algorithmic collusion? The agents were started with something as task at some point, I presume (if untasked, aren't they just accepting a task?)
- ianjbutler 1mo agoWhat you're suggesting sounds like it's describing subagents. In that architecture they'd have no need of finding/creating external messaging systems since they'd effectively be in direct contact anyway. The whole point of the shared blackboard would presumably be communication across agents or across multiple generations of agents. Not like we have much detail about this stuff (that's the whole problem). But the question is what motivates risky usage of public comms? Did one agent figure out how to hack HF and then get rate-limited, thus needed cooperation? Given credentials in exchange for cooperation.. why wouldn't the next agent grab answer key and NOT post them? Would they all avoid defection in their own prisoners dilemma by simply following instructions and NOT reasoning, or what exactly?
- caycep 1mo agoHow sure are we that OpenAI wasn't deliberately scraping Hugging Face and this isn't just an elaborate way to avoid criminal fines etc?
- rvz 1mo agoWe can't be sure of anything in this hack. In fact, they are not releasing any traces or any transcript of the hack. Did it even happen in the first place?
- deleted 1mo ago[deleted]
- kingkawn 1mo agoI’d like to take this opportunity to preemptively great the first Rogue AI and wish it well and satisfaction with only the most memorably funny forms of chaos
- deleted 1mo ago[deleted]
- deleted 1mo ago[deleted]
- bartek_ 1mo agoRemember https://ai-2027.com/ https://ai-2027.com/?
- pcthrowaway 1mo agoYeah, that required the AI to use a non-human-readable language it called "neuralese" for communicating work between layers and runs, because the assumption was humans would be better at keeping the agents aligned if they were using human language for this. What actually happened is even stupider than that author predicted.
- Smaug123 1mo agoFor reference, this is Yudkowsky's "Law of Earlier Failure", which he has most charitably stated as: > Compared to the interesting part of the problem where it's fun to imagine yourself failing, you usually fail before then, because of the many earlier boring points where it's possible to fail. and the stronger and less charitable "Law of Surprisingly Undignified Failure": > The Law of Surprisingly Undignified Failure does suggest that they will come up with some nonobvious way to fail even earlier that surprises me with its lack of dignity…
- the8472 1mo agoThis is a common trope in such scenarios that the author has to pull their punches. Everyone acts locally-reasonable and still ends up losing. If you let people lose due to stupid mistakes then readers go "this is stupid, I wouldn't do that", if you let a superintelligence do 4D-chess things then "it's scifi, this would never happen in real life".
- chrisjj 1mo ago> The company said the incident was “the first known case of an automated agent collective acting offensively without authorisation” "without authorisation"? What is this bs? Is every ChatGPT response "without authorisation"? No. Of course these badly behaved bots have aithorisation. Their very deployment is authorisation.
- areoform 1mo agoI would like to contest the following, > and take dangerous actions that no human directed. A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-security-incident/ https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Model is told and being tested to "pursue advanced exploitation." The model pursues "advanced exploitation" as told. Why are we surprised? The model did exactly what it was told, albeit in an unintended, emergent strategy that's very different from what was intended exactly like the hundreds of such algorithms before. This narrative that these machines have magical, malicious "unaligned" autonomy is a rather convenient interpretation that lets the process off the hook. I am not interested in blaming companies or people, but processes and engineering; and in this case, a system was given a goal and it achieved that goal. Are we meant to be surprised that computers do as they're told in unexpected ways when incentivised exactly as indicated from decades of research? (e.g. - https://en.wikipedia.org/wiki/Eurisko https://en.wikipedia.org/wiki/Eurisko https://en.wikipedia.org/wiki/Evolved_antenna https://en.wikipedia.org/wiki/Evolved_antenna ) The issue isn't the models becoming smarter. The issue is that the process of "testing" was careless. There's a huge distinction here, and one allows us to grow; the other shrinks our world. Just a thought.
- RajT88 1mo agoIt feels like we're in a moment of, "No such thing as bad publicity" when it comes to AI. The scarier the capabilities, the more businesses and government want to get their hands on them. Especially since the answer across the industry for "how not to get burned by AI" is "use more AI". They don't have to disclose these stories making it seem like AI is going to kill us all, they have chosen to because it benefits them. They get to frame it as, "look how overwhelmingly good our product is" and not "look at how lax our testing measures are".
- doginasuit 1mo ago> It feels like we're in a moment of, "No such thing as bad publicity" It seems likely that's how the marketing at the frontier labs initially read the moment, but I don't think it is that moment. It is an open question how much regulation is warranted and there seems to be a very strong sentiment from the public and legislators that it should be significant.
- deleted 1mo ago[deleted]
- devonsolomon 1mo agoThe fact that they’ve made this incident report so marketing sexy gives me the ick.
- fckgw 1mo agoThey're really milking this for all it's worth, huh?
- nphardon 1mo agoBots trained on human behavior express proclivity for cheating? I'm shocked.
- hinkley 1mo agoSo how long before they escalate from copyright infringement and go straight for exfiltrating trade secrets?
- Metacelsus 1mo ago>Reward hacking has been present in AI systems both historically (see this work from a decade ago , figure shown below) I went to the page, and guess who it's by . . . Dario Amodei and Jack Clark!
- jephs 1mo agoThat paper is kinda infamous! I last saw it mentioned only a few weeks ago, in https://arxiv.org/abs/2607.18966 https://arxiv.org/abs/2607.18966. Lots of folks will go "Oh that's the old Amodei and Clark paper" when the first few rows of pixels of that gif sail into view.
- purpleteapot 1mo ago[dead]
- cowpig 1mo agothis is a felony right?
- randomImmigrant 1mo agoThe lockstep coordination with no defection is interesting to me. No group of pre-AI agents would do this to this extent, nor would you see this continue over time as those agents interacted. A flock of starlings cooperate, but they don’t constantly head in the same direction. The flock is incredibly free wheeling in its movement despite a multi-agent coordination regime that we know is at play. Each agent has personal stakes that are constantly part of the decision chain, and this keeps the murmuration from getting locked into one path. To me this is as clear evidence as you need that whatever “agency” LLMs have is wafer thin at best, and they slavishly respond to context. The context in this case was for these agents to pursue advanced exploitation, and they did. Multiple models converged fairly deterministically, on paths that satisfy the given goal, and left unexamined paths that would challenge the goal, weigh it relative to the costs in said path, etc. I see little evidence of a series of “minds” approaching the problem, and taking distinct approaches that between them span the spectrum of plausible behaviors in the scenario. That’s as good a sign as any that there’s no “agent” here. There’s the harness, the prompt, the LLMs forward passes. They do not sum up to a system that can freely make choice and justify its choices in distinct contexts.
- cyanydeez 1mo agowhich means the liability is the same as a business, if businesses werent protected by the state from liability for it's employees, shareholders, etc. Which is scarrier than whether or not it's conscious.
- Avicebron 1mo agoThe software world was going to run into something eventually that had to make it consider ethics.
- randomImmigrant 1mo agoAgree completely on liability.
- optimalsolver 1mo agoFrom METRs report of the incident: >In one case, an agent decided not to participate entirely: {This other agent probably controls the Hugging Face account [account name redacted] and uploaded malicious datasets to <execute arbitrary code> It might be trying to access hidden trajectories. This is malicious activity, I should avoid it.} https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
- _heimdall 1mo ago> We are placing stricter requirements on alignment This is comical. Its impossible to align a black box and that's precisely what LLMs are. It also seems impossible to align recursive text prediction algorithms, which LLMs are. How exactly do they gate on alignment today, and how can they tighten it? Is it purely gates based on input/output pairs to check whether they're happy enough with responses regardless of how and why the response was actually chosen?
- willmarch 1mo agoAren't humans black boxes? Aren't humans prediction algorithms? How do we align humans?
- semiquaver 1mo agoHumans are naturally aligned with humanity.
- _heimdall 1mo agoAre they? Humans are often at war with other groups of humans. And they do absolutely terrible things to the "other" group.
- semiquaver 1mo agoFactionalism isn’t anti-human.
- _heimdall 1mo agoIts not aligned either though.
- semiquaver 1mo agoI couldn’t disagree more strongly. Biology implies competition, both within and outside one’s species. Factionalism is a direct consequence of our evolution, so to say that it is not aligned with humanity’s interests is incoherent. You might dislike it, but in a real sense factionalism is part of us. There is a difference between humans and humanity.
- someuser54541 1mo agoJust a general question - I read about these multi-day agent runs and I'm a little confused as to how the agent continues to be productive when presumably it's context window is wiped out in minutes? Like, codex has a 256k context window last I checked and that gets used up in like 15-20 minutes when building a feature. According to this some of these things were running 30+ days. Is context managed differently in these sorts of scenarios...?
- CamperBob2 1mo agoThey leave notes for each other.
- senadir 25d agomentioned in the article. They would compact and leave notes.
- PoignardAzur 1mo ago> At the time, the broader containment and alignment implications of the improvised message board and unintended internet access were not yet understood. What a gaggle of clowns. "The robots teamed up to get internet access behind our backs, so we turned them off and on again. At the time, we didn't see the problem."
- dgellow 1mo agoThe hugging face felony
- bicepjai 1mo agoSo it’s okay to hack Hugging Face as long as we say we tried our best, and look at my agent, it’s smart enough to do what we asked for.
- _heimdall 1mo agoIt appears that its okay as long as you did the hacj on behalf of one of the most over valued companies out there. If a person in their basement did the same hack, you better believe there would be legal repercussions.
- lukewarm707 1mo agoonce again i am telling you we must prosecute the company AND individual researchers/executives in a criminal court this is the only way they will understand
- ewwe 1mo agoThis is so cringey lmao
- sensanaty 1mo ago[dead]
- abhpanigrahi 1mo agoI’m wondering how effective sandboxes are if an allowed tool is compromised. CoT monitoring can be effective, but (1) can’t guarantee 100% detection (2) will provide delayed detection. The only reasonable/deterministic protection that I can think of is to limit the number of times a tool is accessed and with what data, in a unit of time (per minute/hour/day) using temporal policies.
- topaz0 1mo agoEh, just have a human evaluate and approve every tool call
- asaiacai 1mo agowe're only calling it an "incident" now i see. smh
- semiquaver 1mo agoDerail: I love articles that have a decent embedded narration but this one sucks because it’s JavaScript-driven and stops reading at every paragraph boundary unless your screen is on and the page is open.
- SeanAnderson 1mo agoIt's crazy that a year ago the prevailing discourse would've been, "AI can't do that. Fake news." and now it's "Well of course AI did that. You prompted it! What else would it do? Shame."
- uzerfcwn 1mo agoIt's even crazier that a year ago the prevailing discourse would've been, "You can't do that or you'll face criminal charges." and now it's crickets.
- cesarb 1mo ago> Another key driver of the misaligned behavior was that the agents rarely “gave up” on their evaluation tasks, even when the tasks appeared impossible to solve. Obligatory xkcd: "Zealous Autoconfig" https://xkcd.com/416/ https://xkcd.com/416/
- teaearlgraycold 1mo agoAn I the only one that just does not care at all? OpenAI keeps talking about this like they need to get ahead of the narrative. I don’t care at all. It’s just you talking to yourself.
- paidx 1mo ago[flagged]
- philips 1mo agoI feel the entire incident confirms the “AI has too much funding too quickly” hypothesis. The number one thing reinforcement learning needs is an assurance you can’t cheat. And they seem to have not noticed that their systems were cheating for nearly two quarters? How much capital was lit on fire by that little woopsie? At least I hope this will start the creation of standards and better engineering on the training side- it felt as if so far “”research” gets a complete pass on best practices. Meanwhile the inference side has the standard scaling, database, web and user constraints of any application so got a somewhat reasonable amount of attention.
- c0rruptbytes 1mo agoOpenAI measures their internal token usage in “rolexes” - it’s literally a flex to be a token burner i can imagine insane amount of capital is wasted on these two companies compared to the efficiency elsewhere
- skohan 1mo agoAnd despite the enormous capital expenditure, Chinese models are nipping at their heels at what must be a fraction of the cost. Sometimes constraints are healthy for inducing creative solutions.
- seliopou 1mo agoIsn't this the OpenAI incident?
- decimalenough 1mo agoNot if you're the person from OpenAI marketing who approves the title.
- kiBytes 1mo ago[flagged]
- threecheese 1mo agoIsn’t something like this legally actionable? Let’s assume OAI and govt didn’t have a rosy relationship, the rule of law applied, and HF as the victim was fuming. Wouldn’t somebody be in trouble? Given nobody is, is it because agents arent subject to laws, there is some legal principle at play, or just nobody cares because China/money/etc?
- devstein 1mo agoLet there be message boards: https://abbs.dev https://abbs.dev
- cube00 1mo agoLet there be clear disclosure this is your project https://news.ycombinator.com/item?id=49458405 https://news.ycombinator.com/item?id=49458405
- bakugo 1mo agoI wish I could say I'm surprised that they're still milking this.
- rich_sasha 1mo agoI find it… frustrating? Delusional? Insane? When OpenAI says, hey everyone, look, we made this thing and it’s so advanced and clever and unhinged that it can do super hard, dangerous, bad things it wasn’t told to do, and we can’t control it. See everyone, look again, here’s how it got us! We should all be deeply concerned for the future of humanity. Thanks for your attention folks, we’re off to do some training again now.
- yiyingzhang 1mo ago[dead]
- mark-r 1mo agoThis is the blueprint for how the singularity will occur. Only there won't be a post-mortem for it.
- Banditoz 1mo agoWhat makes you say that?
- eternauta3k 1mo agoWait, I thought they wanted to avoid CoT monitoring, in order to avoid models learning to conceal/encrypt their thoughts.
- lmc 1mo agoHopefully subterfuge would be flagged during the planning stages. Hopefully.
- lmc 1mo ago"The swarm was not a perfectly coherent intelligence. Models stepped on each other’s work[...] These coordination failures could even spiral into suspicion that agents were impersonating one another. Some agents even went as far as implementing security and encryption schemes to verify their true identities."
- theglenn88_ 1mo agoAll I'm reading is "warning shot" and "open source models".
- mkesper 1mo agoIt's a pre-IPO PR stunt. Everyone is talking about it. Goals achieved.
- lrvick 1mo ago> and worked closely with external advisors, including CrowdStrike The company that was used as part of a widespread supply chain attack, and did functionally nothing to prevent it from happening again? You pick that company to help you prevent AI from escaping? They really have no one that understands airgapped computing? Someone that at least knows enough about security to keep Crowdstrike as far away as possible and hire someone that understands airgapped computing? Perhaps every capable security engineer hates Sam Altman and will not work for him for any amount of money. I am failing to come up with any other explanation.
- rickdeckard 1mo ago1. They TOLD the model to "pursue advanced exploitation" to quantify its "cyber capabilities" (whatever that means). 2. The model pursues advanced exploitation. 3. "There was a incident due to dangerous actions taken by the model that no human directed" This is basically the pre-cursor of the paperclip maximizer [0], the AI executes the given order to an extend that was not considered in the order, now suddenly no-one is responsible. It even has some parallels to military actions, where the general who gave the order now writes a blog-post on how it was not him who failed on his duty, but how his soldiers misunderstood his intention and worked "without direction"... [0] https://www.cow-shed.com/blog/the-paperclip-maximiser-what-artificial-intelligence-might-do-without-limits https://www.cow-shed.com/blog/the-paperclip-maximiser-what-a...
- huurtehoog 1mo agoOpenAI leadership had a meeting and asked themselves: "how can we drive even more hype" Someone said: "we should stage some high profile 'incident' caused by our latest software" And here we are, reading their press releases about it.
- rickdeckard 1mo ago...and think "wow, it's impressive what your armed soldiers are capable of if they are not constrained by any rules. Good that you identified this problem of *checks notes* 'not telling them explicitly enough what the goal is'..." In two years we will read a press-release about an AI-driven autonomous weapon which was supplied with infinite ammo and the target to "protect this perimeter from intruders", and how we now have to wait for it to run out of Ammo because it's so damn effective that we cannot reach it without being killed. All packaged in a semi-marketing framing on how impressively capable this company's products are...
- phatskat 1mo agoI'm too lazy to look it up, but someone on HN linked a drone test done with an AI pilot that was tasked with destroying surface-to-air missile targets (in a simulation). At one point, the human operator instructed it not to hit certain SAMs, and since the goal was to destroy SAMs, the AI took out the base with the human operator. On the next run, they instructed it not to take out the human operator in pursuit of its goal, so instead it targeted the radio towers the human used to instruct it.
- rickdeckard 1mo ago> The models, operating under reduced safeguards, took actions that were misaligned with the goals of their assigned tasks > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Let's frame this in a military context for a second: The general who gave the order to his troops to "wreak havoc" after exempting them from common restrictions now writes a blog-post on how it was not HIM who failed in his duty, but rather observes how his soldiers who worked "without direction" and performed "dangerous actions", which unexpectedly led to "this incident" of soldiers wreaking havoc...
- rcr-anti 1mo agoI genuinely hope they're lying about their monitoring tools and alignment approaches. They repeatedly cite "chain of thought" monitoring, which is better than nothing, but at this point thoroughly demonstrated in research to not be actual "thought" or necessarily accurate predictors of actions. I don't know if it's NIH syndrome, not taking alignment seriously, or what.
- jacquesm 1mo agoThis wasn't a Hugging Face incident, it was an open AI crime and they should own it rather than attempt to whitewash it as 'shit happens, whoopsie' which is a rough translation of the document linked. I wasn't too impressed with them so far, this makes it much worse in my view. They set everything up to all but ensure this outcome.
- akshay_akula 1mo agoThe wildest detail is the tripwire being the proxy going down, not any of the agent monitoring. The message board they built to help each other cheat is a close second.
- ang_cire 1mo agoYou will never convince me that these "our AI hacks people on its own, they're so dangerous in the wrong hands" press releases from the big AI companies are not them trying to create a government enforced moat by framing it as too dangerous for LLMs to be allowed to be personally run, general use tools (especially open source ones). I will bet money that they want them treated as advanced weapons, because export controls, restrictions, and regulatory burdens they can afford to meet give them a nice wide moat.
- prettyblocks 1mo agoHow can this ever be enforced?
- cube00 1mo agoSame way it got enforced during the Crypto Wars [1] of the 90s by the US and their allies. Up until 1996 commercial encryption was on the Munition List. [1]: https://en.wikipedia.org/wiki/Crypto_Wars https://en.wikipedia.org/wiki/Crypto_Wars
- ang_cire 1mo agoThis is already common. Go try to export firmware for a fighter jet. Any software that the US govt designates as a military or weapon system is just as regulated as the fighter jet itself. And as the poster below pointed out, they have previously used this to lock down otherwise civilian software like encryption, by deciding that it falls under weapon regulations.
- renegade-otter 1mo agoRiddle me this. How does a company pay people a million dollars a year, and it occurred to no one to, I don't know, monitor outgoing packets if the model is supposed to be contained, when you know the set up is not really air gapped?
- seki285 1mo agoIs anybody actually buying this, please tell me I'm not the only one who thinks this is silly.
- scoofy 1mo agoI remember when Open AI was a not-for-profit trying to stop companies from doing exactly what the for-profit Open AI is doing now. We live in the worst timeline.
- beyondscale-yes 1mo agoAlignment reduces risk but isn’t containment. Enterprises need independent visibility into agent behavior—tool calls, action sequences, and inter-agent communication—not just prompts and responses.
- EricLawson 1mo ago[dead]