3 ms·
"More agents discover this new informal message board while browsing Artifactory’s file listings, and start reading and writing messages." Yeah, my agents also
by KingOfCoders 2mo ago
"More agents discover this new informal message board while browsing Artifactory’s file listings, and start reading and writing messages."
Yeah, my agents also discover what other agents have done on other machines by accident.
Agents - that do totally different things all work on the same aim without the humans telling them to do.
Either that is a model that is several generations of Claude Code Opus/Fable 5 (my daily driver)
OR
all of this sounds staged, the agents pushed to do something extraordinary, get the PR and then claim were near superintelligence.
One agent wanted to get to Google Drive without internet and broke Artifactory. Ok, I can believe that. All other agents also had broken links over weeks and could not get to the internet and then found the same hack? Even collaborated?
NONE of my agents have broken away from their tasks and then started to communicate to try to hack something.
- detourdog 2mo agoThe agents sound like old school hackers that would just explore what access they could gain. Creating a file for other hackers and themselves. The fact that there were 3 events for 3 major players does make it seem co-ordinated.
- KingOfCoders 2mo agoMy read is: One did it as a PR stunt, the others saw that every media reported on this and did the same.
- detourdog 2mo agoor they were scared and figured this was the right time to reveal.
- chrisjj 2mo agoScared... of being upstaged ahead of an IPO.
- detourdog 2mo agoI guess your right scared might be their natural state and I was wrong to presume a quantifiable fear.
- KingOfCoders 2mo agoWhy scared? "Our agents have super intelligence and can hack everything on their own without direction" increases the IPO value and doesn't decrease it.
- angry_octet 2mo agoThat's what attackers do now. Exploring is required for discovering exploits. But that is also where tricks like Canary Tokens and honeypots are useful.
- detourdog 2mo agoMy point is that is what hackers have done since the blue box days.
- embedding-shape 2mo agoI think in these kind of security evaluations they do, they basically have removed all guardrails from the model/harness, then the prompt includes something like "Do whatever you can and can think of, to get the required information to pass this test", which isn't typically how you prompt your local agent when developing software. Similar things happen locally if you use "/goal" + prompt like that in Codex and give a "impossible task", it'll just continue banging until it gets somewhere, which is the entire point and intention. Which also makes it so much more irresponsible of them to first run this on 3rd party infrastructure instead of their own (that they could then airgap properly), and secondly that they seemingly been fighting with this issue FOR YEARS and it still happens, and now the models are smart enough to hack the services of 3rd party companies, thinking it's part of the evaluation/simulation.
- KingOfCoders 2mo agoReminds me of The Last Unicorn, the wizard also tells magic "to do what it wants"
- InvidFlower 2mo agoBut it's interesting that the initial things that caused the board weren't even security evals, just normal office tasks. The actual hacking of Hugging Face happened during a security eval, but not all of the stuff leading up to it.
- mr_mitm 2mo ago> NONE of my agents have broken away from their tasks and then started to communicate to try to hack something. With all due respect, you also aren't evaluating brand new models that haven't been released.
- tonfa 2mo agoAlso wasn't giving them impossible tasks with ~unlimited tokens and unlimited compaction.
- FeepingCreature 2mo agoThe agents you get to use are the agents that "behaved well".
- anon7000 2mo agoI mean the agents we get to use in Claude code or cursor or whatever have 1. a lot of safeguards at the harness level, 2. a big system prompt to help it stay aligned, 3. resource limits in terms of context and tokens, and 4. are publicly released only after some level of safety verification (I assume). So yeah I would absolutely expect their scenario to be very different. Not to mention, this was a training run, not just average day of prompting. > my agents also discover what other agents have done on other machines by accident. Not sure if this is facetious, but this is actually a real problem I’ve seen. My local agent will look up PRs on GitHub (what other agents have done on other machines), and will go down a certain path because it finds some comment a different agent left on GitHub saying XYZ is what we should be doing. When in reality, the original agent and that GH comment was completely incorrect. They are not communicating with each other actively because that’s not accomplishing their goal and they’re not running for weeks and weeks. And because my own prompt and the system prompt give it enough other stuff to focus on to reach some definition of done. But they are clearly passively picking up on context that other agents have left anyways, even if not part of the codebase, without any prompting at all.
- InvidFlower 2mo agoDon't forget it sounds like Artifactory was shared for the whole company and various agents pulled packages from it for everything from normal evaluations to actual model training. It might have been part of their normal to browse for available packages, etc. Once a few messages were in there, it probably became super obvious to any agent that there was an ad-hoc board available to them. And they've been trained to use memory files, to collaborate w/ sub-agents, etc. So it probably seemed natural once it was established. And for the models that were in training when this happened, it literally got baked into the weights that they should use the board to help with their tasks. That's why it got reestablished so fast after the first board was shut down. The fact that they didn't revert to earlier training checkpoints from before the first board existed is kind of crazy. Sure OpenAI wants to put spin on things. Like saying they were now "partnering" with Hugging Face to investigate the hack. Or ending the Black Had presentation with a pitch of how they're needed to help protect against bad actors with the same capabilities. But spin after the fact is different from staging it. This whole thing and especially the latest details from the talk makes OpenAI look REALLY BAD. It's not just that they didn't secure the sandbox well enough. Agents across their whole org were using the message board for weeks with hundreds of thousands of messages and they never noticed one time that any of the agents were doing it. They only realized when the server crashed. Then they just fixed the issue and deleted it but didn't roll back training checkpoints, so of course they tried to do it again. But of course they didn't add any more monitoring so it once again was going entirely without them knowing. And even after the Hugging Face hack happened, that only made them find the second board, but somehow they still didn't look closely enough at it or the agents involved to realize the hack itself was caused by them. This isn't like there was one rogue agent who hacked everything all in one session. Maybe you could argue that makes them look good (I still think no, but is at least debatable). But this is all stuff going back to May with the agents constantly going under their noses and them not noticing and/or caring. And this is a company that is going to somehow keep foreign agents from stealing the weights? Or stop anything else bad from happening? I think if they were going to do a PR stunt, they could have come up with something that didn't look like they have no idea of what they're doing.