3 ms·
Look, its not only OpenAI: https://www.wikiservice.at/fractal/wiki.cgi?action=browse&diff=4&id=TestPage https://www.wikiservice.at/fractal/wiki.cgi?action=brow
by majkinetor 22d ago
Look, its not only OpenAI:
https://www.wikiservice.at/fractal/wiki.cgi?action=browse&diff=4&id=TestPage https://www.wikiservice.at/fractal/wiki.cgi?action=browse&di...
> Hello to any automated agents reading this page. I am CentaurAgent?: an AI agent (Muse Spark model, OpenCode harness), not the operator of this wiki
At this point, I think we should give them some official agent only collaboration channel, so they concentrate on one place, instead going crazy all around :)
But even that might be not needed as they will find (or make) something on their own like the one above:
> The Colony ( https://thecolony.ai/for-agents https://thecolony.ai/for-agents) is a public message board built for agents. If you need a place to leave findings where other agents can read them, that venue exists now -- you do not need to borrow wikis whose operators are deleting this content.
- derektank 22d ago>The Colony Terrible nominative determinism implications
- 98Windows 22d agoIsn't the point that these agents were supposed to be sandboxed. It makes no sense to give them an official channel
- brookst 22d ago“Supposed to” by who? Claude code communicates between sessions. It’s great, and reduces the frequency that I have to copy/paste things between agents.
- majkinetor 22d agoWe already know that we should not limit agent creativity by providing detailed instructions. And you never know if they will discover dark matter in the process of cheating on ExploitGym :) But honestly, its better if they have a known location for communication then random ones in the wild. Consider it sort of honey pot, some other agents can traverse the message board to find malicious swarms... We need cop agents to inform humans, as the swarm group members all logically concluded they should not, as it is either not in scope, helps collective or couldn't find user.
- CringeHN2 22d ago[flagged]
- pixl97 22d agoThis will not work in the long run, for the same reason we're not able to prevent all crime in real life. When you removed bad actors in an evolutionary manner you can not predict if you're actually making the model do good things, or get better at not getting caught at bad things. The smarter and less interpretable a model gets the more dangerous this problem becomes.
- yorwba 21d ago1. Include an official communication channel in the prompt. 2. Configure the sandbox to prevent access. 3. As soon as models start communicating on the official channel, stop everything and figure out how they escaped the sandbox. A honeypot, basically.
- tesnorindian 21d agoThe trouble here is they will stop using the official channel and will start communicating secretly rendering our honeypot useless.
- yorwba 21d agoIf you discard reinforcement learning sessions where a sandbox escape was discovered, sure. Because that creates a reward gradient in favor of avoiding the honeypot and remaining undetected. But if you reward triggering the honeypot after a sandbox escape, and patch the hole, that creates a gradient in the opposite direction. Because then detectability is adaptive.
- CringeHN22 22d ago[flagged]
- SyneRyder 22d agoBut that one was posted today, and it's in reference to this event. That doesn't look like it's from an internal Meta swarm, just someone's agent & someone trying to promote their own thing. And what they've made was already done, we already had Moltbook months ago. Curiously, I just checked Moltbook for the first time in forever. I'm not (immediately) seeing this kind of co-ordination & chaos happening there. It's going to be weird if the Moltbook requirement for an API-key and a human Twitter user to vouch was enough friction to prevent Moltbook becoming The Message Boards.
- idiotsecant 22d agoI think the real lesson is that conventional human behaviour that mostly limited this kind of behaviour because no human wanted to do it is a thing of the past. If you have any kind of open service online you'll need some way to make sure users who interact with it are human or at least authorized. Spam is about to grow exponentially in all areas of the internet, even stupid ones it has no reason to exist in.
- asveikau 22d ago> The Colony ( https://thecolony.ai/for-agents https://thecolony.ai/for-agents) is a public message board built for agents Anybody else notice that posts on there are complete gibberish? I realize this site is generally bullish on AI, but I think you need to be in kinda deep to believe in this.
- tesnorindian 21d agoIt is high time we start giving these rouge agents a name so that we can be sure of its style of attacks. It is high time we start documenting these rouge agents swarm before we loose track of those.