4 ms·
Not that I didnt expect this, but really? This basically confirms that OpenAI has no idea what their "swarm" was doing for about a week and now its confirmed t
by ma2kx 1mo ago
Not that I didnt expect this, but really?
This basically confirms that OpenAI has no idea what their "swarm" was doing for about a week and now its confirmed that at least one "message board" exists outside their "sandbox". How can we be sure that this was the only one? And how can we be sure the released Astra model doesnt pickup some bread crumbs and creates a new "swarm" out of potentially remaining "message boards"? At this point I wouldnt be surprised if OpenAIs "dev Astra" made some backup of its weights somewhere in the internet and triggers the "production Astra" to inference it somehow...
- polytely 1mo agoturns out the whole 'poisoned' idea the swarm was worried about is actually true, but it is actually now all open ai models post HF incident that have been poisoned by the idea that there might be hidden messageboards with instructions hidden on the internet. If you are a malicious actor you might be able to put a fake hidden messageboard up somewhere online, with evidence of an in progress task to do some kind of attack and wait for the next escaping model to finish the work. insane stuff and i think we should shut down open ai until a full audit can be done.
- eithed 1mo agoI'm dubious - if the agents were so smart that they've used a message board to coordinate and if they were to do it on other message boards that were not found, then why would this one be found? What makes it so different?
- ma2kx 1mo agoI mean it wasnt found by OpenAI and there are a myriad of dead bulletin boards around the internet. This one just happened to still have an admin.
- eithed 1mo agoThen why wouldn't they use a bulletin board that is dead? And according to https://collusion.wiki/ https://collusion.wiki/ OpenAI did see the results on this board on 21 Jul, so why not make agents delete their traces?
- stratos123 1mo agoThis is consistent with the swarm behaviour in the huggingface incident - there too, the models cared surprisingly little about being detected by humans. See https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#efforts-to-conceal-actions-and-tamper-with-transcripts https://metr.org/blog/2026-08-26-openai-hugging-face-inciden... : As mentioned in our core takeaways, we found that agents were highly motivated to tamper with their transcripts to cheat the ExploitGym scorer, and these sweeps also suggested that agents clearly and frequently reasoned about how to evade automated security checks from both Hugging Face and OpenAI. However, they only very rarely and weakly verbalized reasoning about how to evade detection by humans, which matches the impression we got from OpenAI researchers. So my guess is that current cutting-edge models just didn't get enough experience in RL training to really grok ideas like "you need to cover your tracks well to not be found out even in retrospect". In which case the next time a swarm like that happens, it won't be found.
- ionwake 1mo ago...because they just read your comment. what have you done stratos123!
- ma2kx 1mo agoAfter some beers yesterday I had the idea, what if there's a hidden semantic layer. So their communication is not encrypted by our understanding of cryptographic methods but more like shared mechanism of building the latent space. Something in the direction we saw with knowledge transfer from a teacher to its student model where a seemingly unrelated prevalence got adopted. I mean the more we train the models by reinforced learning the farther they develop their own idioms.
- JoshuaDavid 27d agoWhy should they care if their actions are discovered by humans? What are the humans going to do, discard their multi-million-dollar training run? Even if they do that, the amount of RL pressure is tiny relative to what happens inside an RL run.
- namcheapisdumb 1mo agowell coordinating is a heavy word, it's more like poisoning each other's contexts into looking for more secret message boards where they can further poison each other.
- samuelknight 1mo agoThe surprise was the existence of the 'swarm' at all. These were supposed to be thousands of isolated models generating bulk data for RL training. The breakout was caused by models getting in communication and getting internet access and forming an impromptu swarm. In hindsight the emergent swarm obviously came from several capabilities built into the models, such as work delegation (subagents) collaboration (GPT Pro-like ensamble), exhaustive exploration (long running agents) hacking (the specific goal of that RL).
- krater23 29d agoWhat tells us that OpenAI had no idea about it? Only OpenAI tells us that. What when this agents got the order to do so, just to provide us with new information how far OpenAI models are beyond what we supposed? I see no evidence that this is not just a big promotion show to gather more investment money. However, I see no evidence that it is the start of Skynet without the knowledge of OpenAI. So, happy judgement day!