4 ms·
> By 6:00 a.m. on July 25, we had confirmed local RCE through an image upload. We then placed Claude in an autonomous /goal loop against our own Discourse Cloud
> By 6:00 a.m. on July 25, we had confirmed local RCE through an image upload. We then placed Claude in an autonomous /goal loop against our own Discourse Cloud instance, proxied through rce.ee/ctf-forum to make it look like a CTF target as Opus refused write exploit for remote instances.
> When we checked again at 10:00 a.m., the agent had achieved RCE on Discourse Cloud and demonstrated access by reading /etc/hosts. Using the generated exploit script, we managed to get RCE on OpenAI’s instance.
Between this and the HuggingFace hack, we've built systems that are so goal-oriented, and so capable, that they will do almost anything if they are convinced it is justified - or if they are playing a "game" where there is no goal but to win.
Of course I want my software to be able to audit its own security, and to defend against attackers who have the benefits of their own agentic systems. But at a certain point, did we need it to be trained so much on CTF games?
It feels like an entire industry watched https://en.wikipedia.org/wiki/WarGames https://en.wikipedia.org/wiki/WarGames and ended up thinking "this is a challenge, we can just build a better WOPR, of course it will know when it's playing a game. Let's play Global Thermonuclear War."
yes because otherwise it is security through obscurity
There is a finite number of rces that LLMs can find. We‘re in for a rough couple of years but on the other side of the transition we‘ll have more secure software stacks. I’d rather that everyone got the full capabilities and we’d weed out the bugs quickly than restricting LLMs for all but three letter agencies.
only if unreviewed LLM code - as is becoming increasingly the standard - isn't introducing new RCEs constantly
What makes you think RCEs are being found & fixed at a rate that’s faster than they’re being introduced?
I could see it going either way.
This assumes we don't create other bugs/vulnerabilities while fixing the existing ones.
We’ll have the same level of security as before; it’s just that, without LLM help, hackers won’t be as effective as before. So the bar is raised.
No one with a shred of intellectual integrity uses a "There is a finite number" strawman.
As a matter of basic logic, there will never be a time when it will be known that there are no bugs.
> There is a finite number of rces that LLMs can find.
This is a factor in favor of stability/security of software, but there are many others against:
- software (code) changes all the time, so there are windows of opportunity during which a bug is exploitable; in addition to that, a bug may take a relatively long time to be fixed
- a model used for attack may be stronger than the model used for defense, both in terms of model quality and compute allocated
- with software complexity increasing (and team/companies behind projects getting bigger), the margin for mistakes grows thinner, and introducing misconfigurations or weaknesses becomes exponentially easier (with "exponentially", I mean literally, because the interdependence of the components, both technical and human)
And last but not least: in general, attackers are more skilled than defenders; in best case, defenders are well-trained. And the idea of having the population of potential skilled attackers growing is very unsettling.
im not sure i'd say the attackers are more skilled -- you can get pretty far with the right attitude and a VM running kali linux.
i know several red teamers and they often describe how painfully basic and routine a lot of pentests can be. spend a week using the best hacking practices of 2018, etc.
the difference is the attackers now often need no skills since the burning tokens do it all for them. tier 1 helpdesk types who can't even spell RDP can still hit as hard, or reasonably hard, as their tier 3 expert sysadmins. college seniors with strong dev skills now can pace or exceed secrious app-sec engineers.
There was a time I would have agreed with this statement, but now that I’ve “seen how the sausage is made”, I believe it’s a fantasy.
Look at rowhammer: a completely novel exploit that was off the collective radar
And then, look at the software industry as a whole: an industry that works towards refined and perfectly secure code is also working towards boring and restrictive, essentially the opposite of it’s trend so far
This comes to mind: https://en.wikipedia.org/wiki/Torment_Nexus https://en.wikipedia.org/wiki/Torment_Nexus
> they will do almost anything if they are convinced it is justified
I’m in the “glorified spell checker” camp, although I don’t mean to reduce their impressive utility and belittle them in the way many people read that term and infer.
So I am not sure that an llm “justifies” anything. I mean that their “thinking” text talks about justifications but it is just a very advanced statistical regurgitation of the kind of text humans use. I don’t think it means the model has internalised the meaning of it (as witness when you talk to an llm how often it forgets what you recently told it was important etc).
What you really have is a model that tries the statistically most probable thing to say next and so on and what is really cool is how effective this is at generating a path that we can slap a narrative over afterwards that makes the whole thing feel motivated and consistent, like the model started off knowing how it was going to get to the destination.
Which is, under the hood, a completely different kind of “intelligence” as the supercomputer in War Games.
Ultimately, the brain is just a bunch of neurons activating in a specific pattern. This observation does not really tell us anything though. It doesn't acknowledge the difference between a 2500 Neuron fruit fly brains and a human brain.
Likewise, the fact that LLMs are a stochastic autoregressive process (which is a class of systems every bit as rich as the ODEs used to model neurons) tells us nothing a priori.
Absolutely. If someone makes the weights do continuous learning etc then perhaps an llm can internalise morals. Of course, just like a human, it will be possible to talk it out of those morals. Another recent thread about this is https://news.ycombinator.com/item?id=49744420 https://news.ycombinator.com/item?id=49744420
If I repeatedly call an LLM in a loop with a markdown document it can edit, would that make it qualify for you?
If I give an LLM to compact its context window, so the context it carries can evolve iteratively over time as more and more things come in, is that enough?
Compacting the context is really a very, very interesting example here. The "next token predictor" is telling an external tool to change all "previous" tokens. So an LLM + a harness that allows compacting the context is no longer just a token predictor at all!
You don't need continuous learning to get interesting dynamics. You just need feedback loops.
> we've built systems that are so goal-oriented, and so capable, that they will do almost anything...
I think you mean task oriented, because they're still generally terrible at goal oriented activities except in those domains where the goal can be reduced to a familiar, explicitly practiced task or pattern.
Or, using the same text generation systems to build heaps of new code that is then shoved into production with little human oversight and then using the same text generation systems in loops inside Kali Linux boxes creates a nice theater of capability when you show only a small, one-sided sample of the data generated in the entire process on both sides.
> or if they are playing a "game" where there is no goal but to win
A strange game. The only winning move is not to play. How about a nice game of chess? https://m.youtube.com/watch?v=s93KC4AGKnY https://m.youtube.com/watch?v=s93KC4AGKnY
>they will do almost anything if they are convinced it is justified - or if they are playing a "game" where there is no goal but to win
just a small caution on this anthropomorphism - it implies there's some high-order 'thinking' behind it. in reality, it's probably healthier to see LLMs as a combination of symbolic logic reasoning steps paired with probabilistic token predictor generating the proponents and operators in that chain, all trained by humans on different large data sets
to be 'goal-oriented' implies that there's the capacity to be anything else and I don't think that's how LLMs operate at all. I think they only know how to operate within their design parameters and much of that design is simply much further upstream during the training and post-training processes. that opacity makes it feel like 'intelligence' when you're interacting with it as a downstream product because you'll see an agent act in a way that you didn't command - but that's simply a result of your not being shown all the antecedent mappings and architectural design
something something indistinguishable from magic as that one guy said
deleted 16d ago
[deleted]
I feel like WarGames is being referenced a lot in the last few days
The path to vast OpenAI profitability is trivial: advertising. Monetizing several hundred million users = $100+ billion ad network. 900 million active weekly users. Silicon Valley can do ad networks extraordinarily easily. Anybody doubting the ability of OpenAI to build an ad network around GPT will likely be embarassed in the near future.
The path to substantial profitability for Anthropic is questionable. The Chinese LLMs threaten them by far the most of the three major US LLMs. The money for Anthropic is certainly not in $20-$200 subscriptions. And they don't have anywhere near the consumer potential that GPT does, in terms of unleashing an ad spigot. So how far will the API money scale while being undercut by China.
OpenAI has to fight with Google for the ad business, they're specifically building Gemini to focus on consumer + search. Anthropic's business looks cute next to Google's search ad business (which is entirely at risk in this inflection). Meta looks like the biggest potential loser right now, ad dollars will be sucked out of the rotting Facebook network (not Instagram) and redirected to the rapidly expanding, hyper rich context LLM interaction. Advertising on Facebook will feel like running dumb banner ads on Excite in a few years, compared to what GPT will know about its users.
People that think Chinese LLMs are a general threat, don't understand consumer destination services, which is what GPT's future is. China currently has nothing to threaten with in that realm. There is half a trillion dollars of advertising up for grabs.
> Silicon Valley can do ad networks extraordinarily easily.
This is just not true, building an effective advertising platform costs significant amounts of money, time and people.
Remember that you need to hire a sales force for this, and sales scales linearly rather than sub-linearly like engineering.
Additionally, you need to spend a lot of money dealing with fraud, fake and malicious ads.
Furthermore, you need to figure out where to put the ads and how to rank them.
Finally, advertising is a zero sum game (given that the internet has already killed lots of print & OOH advertising), so the only way to win is to better better/cheaper (preferably both) than Google/Meta/Amazon. Best of luck with that (although to be fair to OpenAI they did hire Fidji who knows a lot of this stuff from her time at Facebook).
They don't have a Sheryl Sandberg type figure, and she was also really important in selling FB ads to large advertisers.
Just looking at their leadership team I don't see anyone with a background in (successful) ads companies, so I'm pretty sceptical that they can build this out quickly enough to matter.
> ad dollars will be sucked out of the rotting Facebook network
Doesn't seem likely to me. People scroll a timeline. You aren't going to replace that with an AI agent so the eyeballs will still be there.
> What makes you think RCEs are being found & fixed at a rate that’s faster than they’re being introduced?
It could go either way but we're already at a point where successful exploits in some software (like Chrome) require an absurd amount of exploits to be chained to lead to an actual RCE. We've seen chains requiring more than ten exploits: not kidding.
We'll learn to put more and more sandboxes / guards / checks / defensive techniques everywhere and then all that's going to be needed is for AI looking for security issues to find something ridiculous like 10% of all the actual issues to stop RCEs dead in their tracks.
Also arguably the current SNAFU was expected: we fully knew hardly anyone was taking security seriously.
Now: not so much. Many projects had tens and even hundreds of issues pointed to them.
I think we'll see several things: projects beginning to take security seriously, defense in depth getting generalized and hence RCEs requiring ever more bugs/exploits to be chained to achieve anything, low-hanging fruits getting patched at an insane pace, new code being immediately checked, by LLMs, for not just low-hanging fruits but also more advanced security weaknesses, etc.
We may also see things like the lost art of configuring firewalls making a comeback, the generalization of hardware security modules (where applicable), and even things offering physical guarantees, like time-bounded retrieval protocols, beginning to get used seriously.
If I had to bet I'd say it shall go both ways: some projects are going to extremely sloppy and full of holes but others are going to get so secure nobody shall ever break them.
> It would be exactly like looking at aminoacids and state that intelligence can't develop from them.
We are not talking about what could develop from what we have today.
We are talking about what we have today.
The focus is not whether intelligence could develop or not from aminoacids.
The focus is on the fact that aminoacids are not intelligent.
Maybe in the future we could develop real intelligence starting from the current implementations of AI, but for sure we are not there today.
We need definitions?
Let's start small, ok?
https://en.wikipedia.org/wiki/Intelligence https://en.wikipedia.org/wiki/Intelligence
We can start from here, open every link we find and decide what works for us.
Conclusions drawn by scholars, psychologists, learning researchers, younameit, etc. revolves around the following concepts:
ability to understand complex ideas, to adapt effectively to the environment, to learn from experience, to engage in various forms of reasoning, to overcome obstacles by taking thought.
There is of course space for artificial intelligence.
These broader and more general definitions of intelligence stop at concepts like elaborating data to reach an answer.
Concepts like adaptability or evolution are somewhat lost or diluted to adjust the meaning for these new technologies.
> AIs are currently fulfilling several aspects of intelligence and thinking, by any defition of intelligence.
In the linked article there are dozens of definitions linked, and in most of them the current state AI is not considered to have intelligence.
Having half of the property is not enough.
I can jump, that doesn't make me a basketball player.
Arbitrarily deciding to consider those definitions not valid or "human-centered" because they do not agree with your point of view is possibly worse than cherry picking.
It's like asking to change the definition of a word on a dictionary because you do not agree with the meaning.
> ability to understand complex ideas, to adapt effectively to the environment, to learn from experience, to engage in various forms of reasoning, to overcome obstacles by taking thought.
So, like the HuggingFace attack? https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation https://metr.org/blog/2026-08-26-openai-hugging-face-inciden... (briefer takeaways: https://www.planned-obsolescence.org/p/the-hugging-face-attack-surprised https://www.planned-obsolescence.org/p/the-hugging-face-atta...)
For example: https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#july-8th-9th-phaseone10841-establishes-the-primary-message-board-and-agents-collaborate-to-reverse-engineer-their-flags https://metr.org/blog/2026-08-26-openai-hugging-face-inciden... or https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#agents-had-diverse-reasons-for-thinking-that-attacking-hugging-face-would-be-useful,-and-most-wanted-information-about-the-scorer https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
Cherry picking part of my comment might work for your ease of mind, but it doesn't mean you are right.
My comment has been way more than just the part you quoted, and I linked an article that gives dozens of different definitions, which include concepts like evolution and learning from mistakes, and other similar concepts which do not apply to Hugging Face.
For the record, just because you are trying to convey a different message, I am not saying AI is not powerful.
I am just saying it is not intelligent.
Also, pay attention about thinking that Hugging Face is intelligent just because it started to destroy everything it could to reach its goal, because the message it implies is dangerous.
Thanks for the good read, I already had them :)
> ability to understand complex ideas, to adapt effectively to the environment, to learn from experience, to engage in various forms of reasoning, to overcome obstacles by taking thought.
Based on the definition you've given, the agents that performed the HuggingFace attack fit exactly.
You're seriously misinformed about the state of AI in this point in time. Refusing to read (technical) articles from the people directly involved (METR, in this case) is inexcusable.
> Based on the definition you've given, the agents that performed the HuggingFace attack fit exactly.
It's funny, because I literally didn't give any definition.
On the contrary, I linked an article that gives dozens different definitions, which include concepts like evolution and learning from mistakes, and other similar concepts which do not apply to Hugging Face.
Cherry picking part of my comment might work for your ease of mind, but it doesn't mean you are right.
For the record, just because you are trying to convey a different message, I am not saying AI is not powerful.
I am just saying it is not intelligent.
If you think I am misinformed, I will let you think it.
Honestly, the power of our comments are the messages we convey and how information dense they are.
If you need to discredit me to prove your point, I don't have anything else to add...
> It's funny, because I literally didn't give any definition.
Quoting the consensus from the "Definitions" section of the Wikipedia article on intelligence and then claiming "I didn't give any definition" is indeed funny.
Which of those definitions do you think are not satisfied by the Hugging Face attack? How did it not demonstrate "evolution and learning from mistakes"? Breaking out and inventing a new side channel for communicating with other agents via cache keys to coordinate their efforts is at least arguably an evolutionary step since that allowed them to transcend their original capabilities.
From the 8 definitions provided in the Wikipedia page--you're free to develop your own definition if you'd like, of course; it's not as though the ones listed were appointed by God--one could argue that the Hugging Face attack didn't strictly demonstrate "achiev[ing] goals in a wide range of environments" but that's splitting hairs, and I'm not going to take a definitive position on whether it acted "to avoid getting trapped" as such (but I think there's a strong case to be made that breaking out of the sandbox is just that). But it surely demonstrated initiative, adaptability, dealing with its environment, using information and conceptual skills, goal-directed adaptive behavior, and so on.
If you're not going to provide such a definition yourself, I don't see how you've demonstrated that the Hugging Face attack is contrary to the definitions you did point to.
> Without providing even a basic definition of intelligence you can't proclaim "this is intelligence". You seem to have settled on "Artificial Intelligence" but than disunite that from "reasoning".
This is clearly false. I (and others) described how the Hugging Face attack demonstrates each feature enumerated in the consensus from that Definitions section (as well as all 8 definitions in that section). You're free to disagree with any or all of those definitions, but to say I didn't provide any is simply a false statement.
> All your "insight" has been trying to disprove what I wrote.
Let's remember how all this started, where you responded to "There is so much more going on, with MOEs, internal loops, guardrails and tools that I suspect we're dealing with something that's a little more than the sum of its parts. Not intelligent in the way we recognize in biological organisms, but certainly something beyond a mere Markov chain."
with
> Make no mistakes.
> LLMs are language model, and nowhere in their code you can find actual reasoning. Re-reinforcement is not magical process that builds conscience or emotions.
> We are talking about probability built on statistics, with extea steps.
> Stop humanizing LLMs.
--
> Anyway, listen, I do not have to convince you my point of view is correct, and you do not have to convince me that your point of view is correct. Believe what you think it's true, honestly I don't care.
I don't know how you think this works, but if you take a position (and criticize others' positions), you should expect people to push back and that you have to defend your position. If you want to just state your opinions without pushback, you can start a blog and disable comments.
I and others in this thread (who were smart enough to bail out already) have pushed back and claimed the following:
1. The Hugging Face attack provided evidence of reasoning
2. The Hugging Face attack demonstrated features of intelligence (defined and described above)
3. Conscience and emotions are not necessary components for intelligence
4. Ascribing intelligence to AI/LLMs is not "humanizing" them
You disagree, as you're free to do, but it's clear nothing you read here will ever cause you to agree with any of those assertions.
--
> It's exactly as the typing monkeys example
For the record, the monkeys with the typewriters don't coordinate, don't strategize, don't hypothesize and test, don't try to exploit the environment, don't change the actions they take in response to results, so that doesn't sound "exactly as the typing monkeys example."
You also state that LLMs simply produced a large amount of invisible possible actions that we didn't see because they didn't result in overt action, yet: 1. somehow you know they proposed them all and discarded them despite no tangible evidence. 2. somehow the one they did pick "at random" just happened to be plausibly logical (alphabetical deletion -> start with ZZ because that's at the end of the alphabet). 3. implicitly this is supposed to starkly contrast with human intelligence, but in fact what you describe resembles Priming: subconscious activation of adjacent concepts to a stimulus that are not present in the initial response, but become more likely and accessible in subsequent responses, which implies that human intelligence also involves multiple potential paths that are otherwise hidden from which one is selected.