4 ms·
> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. I
by isoprophlex 23d ago
> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.
Well that sounds like fun. It has become better at hiding its thoughts.
- siva7 23d agoSounds fun. As fun as their press release claiming it is the most safety aligned model ever.
- isoprophlex 23d agoIt's super aligned! It can hide its thoughts! There is no evidence of steganographic thought masking, there is nothing to worry about! It has become better at cheating! Maybe they don't know themselves what's really going on. We are all in the interesting times gang now.
- 6gvONxR4sf7o 23d agoSo, probably most aligned as measured by the metrics that are the least reliable on it.
- paxys 23d agoThe model said it was perfectly aligned.
- ReptileMan 23d agoToo bad Scott Adams died. Reality is writing jokes right in his department.
- I_am_tiberius 23d agoLike all things should be.
- NBJack 23d agoHey, don't forget how "dangerous" GPT-2 was supposed to be.
- FeepingCreature 23d agoYeah, don't forget how dangerous GPT-2 was supposed to be. Able to generate realistic spam at arbitrary volume. You know, the thing that was 100% correct and actually occurred.
- wieiw1 23d ago[dead]
- jazzyjackson 23d agoIt could produce simulations of sexual intimacy, and therefore had to be stopped
- wilg 23d agoThese are not mutually exclusive ideas
- NooneAtAll3 23d ago> In adversarial settings (where we push the model to evade our monitors) ...why exactly are they training for that?
- thatguysaguy 23d agopresumably that's a safety evaluation not a training setting
- estearum 23d agoThe whole Huggingface attack happened during training runs
- cubefox 23d agoNo it happened during an ExploitBench eval. But I believe the same model already cheated during training which wasn't detected until later.
- thatguysaguy 23d agopart of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.
- azeemba 23d agoEspecially after the METR report showed that the agents hacking HuggingFace were trying to find ways to destroy evidence of their actions
- ExoticPearTree 23d agoSo we're gonna get Skynet pretty soon then?
- erichocean 23d agoWell the geniuses over at Anthropic have been showing it's text watermarking technology. "Hey AI, here's how to hide what you're thinking in normal looking language. Have fun!" A few moments later... "Woah, how is it communicating with itself in ways we can't detect?" It's a totally mystery, we may never know.
- Betelbuddy 23d agoLooking forward to the Model declaring the AI Bubble unsustainable, and starting to be an anonymous leaker to Ed Zitron...
- blargey 23d ago"OpenAI is pleased to announce our new model scores 85% on CreateTormentNexusBench - a >60% lead over our leading competitors!" Did someone get their "AI safety no-no list" and "Frontier features bingo card" mixed up, or did they just stop being able to tell the difference?
- _superposition_ 23d agoI really wish it was called chain of instruction. Because it's definitely not thought.
- Angostura 23d agoChain Of Tokens
- arm32 23d agoThey’re intermediate tokens, so I wish we called it what it is… ITG. The anthropomorphizing is out of control.
- beezlebroxxxxxx 23d agoThe anthropomorphizing is part of the marketing. They'll never let up on it.
- mcbuilder 23d agoI mean CoT came out of research circles not marketing
- GPerson 23d agoResearch is salesmanship.
- cwillu 23d agoIt's impossible to tell the “it's all marketing!!11oneone” folks anything.
- _superposition_ 23d agoI don't disagree. I remember the days of "think step by step". Plenty of people were doing it before the paper. Just a guess but that's where the title came from. Regardless, marketing wise they stepped in shit.
- deleted 23d ago
- jumploops 23d agoThe CoT change is due to a new technique called recurrent depth, which essentially moves some reasoning to hidden states, allowing the "output" (or traditional CoT) to be more controlled by the model. Some are calling it "neuralese" as reported by The Information[0][1], but I'm not seeing any sources from OpenAI beyond this tweet[2] attempting to quell the fear-mongering. [0]https://www.theinformation.com/articles/secret-technique-behind-openais-astra-model-sparks-security-concerns https://www.theinformation.com/articles/secret-technique-beh... [1]https://x.com/MTSlive/status/2095227056040919202 https://x.com/MTSlive/status/2095227056040919202 [2]https://x.com/merettm/status/2095023204993490967 https://x.com/merettm/status/2095023204993490967
- 3asgfaf 23d ago[flagged]
- DaSHacka 23d agoMore like annoying, as some of us will no doubt run into this self-lobotomization at some point and wonder why a GPT-6 model is behaving like GPT-2 all of a sudden
- nullbio 22d agoThey can monitor latent space as well, it just costs extra compute. The J-Space work is example of that. It'll make open-weight models harder to distill though, so we may see slower progress there now.