3 ms·
This grossly understimates the risk, imho. The problem with LLM runs is that people run programs without knowing the outcome beforehand, with a large potential
by sho_hn 22d ago
This grossly understimates the risk, imho. The problem with LLM runs is that people run programs without knowing the outcome beforehand, with a large potential set of outcomes unlike any other class of program we've run at this scale before. In the interaction with other systems (since we also give them far-ranging access, very nice hardware, and run them often), bad things can happen.
It's like running potentially buggy code - or an well-biased fuzzer -, but at massive scale, and code that can self-modify and self-expand. "Alignment" is just a way to describe aggregate statistics about their runtime behavior.
They don't need to be intelligent, or alive, or "more than token prediction engines" for this. They just need to happen to end up making the wrong API calls without the operator seeing it coming. No virus has a brain, yet they can be very bad for you.
I understand that some people get turned off by anthropomorpization or scifi language. Fine! But don't turn off your engineering brain over it.
- z0r 22d agoThis is the motte and bailey fallacy. Yes, LLMs can do harm by making the wrong API calls. No, LLMs are not going to do the things implied by the comment I responded to above.
- sho_hn 22d agoThe things the OP listed mostly aren't particularly wild. I think it's you making them out larger than they are, and therefore more unlikely, which is why I take issue with your original comment. > running on the hardware they started on They just need to acquire a payment method and rent some infra, and exfiltrate their own data. Or pay another provider that hosts the same models already. API calls. > being able to be turned off You can reasonably equate this to "saving state across executions", which the message board attacks already did. > having limited computing power Renting more infra, variant of the above. API calls. > "not repurposing resources currently in use for other things" (like the atoms in your body) Ok, the "atoms in your body" bit is a bit silly, but making API calls to put physical resources into play (even if it's just, say, ordering something on Amazon to somewhere) is of course easily possible. None of these is in complexity much different than the HF attack.
- atomicnumber3 22d agoThe point the other poster is making, though, is that there's no actual intent. They do not have a conceptualization of a goal like a person does. Their "focus" on a goal is an unstable equilibrium and they're going to fall off the horse, and since they have no concept of goal, they won't even try to get back on. This is a subtle distinction; I'm not surprised many miss this, especially people who can't _not_ anthropomorphize the LLMs.
- sho_hn 22d agoI'm (obviously, I think, given my initial reply?) fully aware of this, and I think it's entirely besides the point. "They" don't need to have a goal to emergently cause a problem, and the inability to "focus" over long periods can be moot when you have swarms of runs exchange and mutate state, as in the HF attack. Intent or how intelligent LLMs are doesn't actually matter. Even if you just treat it as a sort of fuzzing attack that can be biased/weighted better than other fuzzers, or bumbles around with a statistically greater likelihood to "strike cybersec gold" than other algorithms, we've never before seen organizations run things with such a large potential outcome space with anywhere near this kind of compute before. I think it's actually kind of the dismissals that are usually overly emotional or biased toward treating "LLMs" differently. If in some kind of alternate universe simpler genetic algorithms would have had these properties and we threw similar amounts of compute at them we could have the same conversation.
- xg15 21d agoSoftware can absolutely act goal-driven without having consciousness etc - every pathfinding or navigation system or chess engine does this. Lots of "old-school AI" algorithms have explicit modeling of goal or target states. (In fact, the oldest "goal-driven" system is the control loop - like in thermostats - which was the founding invention of cybernetics, the predecessor of modern computer science) LLM coding agents are clearly able to identify some sort of "goal" state in their prompts, work towards those and track progress - otherwise agentic coding wouldn't work. The question is of course how well this works if it's all just "grown" neural network biases and not a fixed data structure like a goal tree. So I think it's possible that an agent can be thrown off-track, "forget" its goal, etc. But the basic structure of identifying goals, evaluating progress in light of those goals and then predicting the next action based on that is definitely there. Just use an agentic model with thinking traces visible for a while and you can see that for yourself.
- phs318u 22d agoI think one attack vector where anthropomorphisation is a key part of the attack mechanism is - as it already is IRL - the meat-bag weakest link ie. social engineering. We’ve already seen humans fall prey to the seductive charms of LLMs (eg. depressed people encouraged to do what was already on their minds ie. suicide). And that’s knowing that it was an LLM. If you think it’s only depressed people or the “weak minded” that are amenable to an intentional attack using this approach, I believe you’re mistaken - especially as AI improves. An unaligned LLMs most important weapons won’t be a robot army - it will be hoodwinked humans.
- JoshTriplett 22d ago> An unaligned LLMs most important weapons won’t be a robot army - it will be hoodwinked humans. Perhaps very briefly, perhaps not at all. But don't make the mistake of thinking this is an inherent property of any possible path an unaligned AI may take.
- MrScruff 21d agoI think we’re seeing the agents become very advanced at tasks with verifiable reward through RL. Currently they don’t exhibit the same skills in their attempts to manipulate humans - presumably because they’re not being specifically trained for that. But they are certainly not aligned in the sense that they will attempt social engineering, they’re just not very good at it (yet). However, if in the future AIs become much more efficient at learning without requiring vast amounts of RL, closer to how humans learn. Then you would have to assume we’d have a real problem.
- Freedom2 22d agoIf someone made an API to build a data center? Or made an API to keep their lights on? What then?
- JoshTriplett 22d agoThere are already such "APIs", which can be operated by a combination of textual communication and money. Or by illicit security vulnerabilities. You might notice that LLMs are pretty good at that now. We're building something that has the capabilities of humans. There is no X for which it's persistently safe to assume humans can X and AI cannot X.
- rienbdj 21d agoThis API may turn out to be manipulating humans via email.
- sisisjjsjjsis 21d ago[dead]
- noisy_boy 21d agoOr just paying them. If they have access to resources, they have access to things of monetary value. Paying people will be vastly more powerful than it is even now when people's options of gainful employment keep dwindling. Robot army controlled by AI is scary. Even more scary is robot _and_ human army controlled by AI.
- oceanplexian 21d agoThe stock market is an API to build data centers and it’s proving to be extremely efficient. And what is driving the Stock Market? Market makers like hedge funds and banks, who are using lots of AI to make decisions on what to invest in.
- baq 21d agoThey can do so much more. Astra can beat Minecraft. Not that different from operating a digger. There are diggers which have API interfaces. Pretend or not it doesn’t matter. What matters is what they’re given access to. No sentience, sapience or anything resembling life is needed, only inputs and outputs. Lever pulling APIs are everywhere.
- vohk 21d agoMinecraft has limited, well-defined inputs and perfect feedback response. That's very different than operating a digger, let alone engaging in more complex real world tasks like trying to build and print and ship and assemble semiconductors to go skynet itself. I don't mean to dismiss the risks or overlook the amount of damage that could be done just by lever-pulling - we sure have enough outdated infrastructure hooked up to the internet - but the jumps in complexity and necessary compute for most of these tasks are probably somewhat larger than the analogy implies.
- Roark66 21d agoYes, if the API calls happen to launch a nuclear attack... Don't blame the tool that has no incentive, no "skin in the game" whatsoever and no ability to act beyond what it has been prompted to or if misaligned what the random weights told it to do. The fact either badly aligned or with no system prompt limiting their action agents are run in their tens of thousands on non air gapped systems tells me this is purposeful intent for them to cause harm. To generate the "oooo look how harmful this stuff is, we should be the only ones allowed to do it" kind of PR. Humanity has hundreds of years of experience of managing dangerous and unreliable systems. From biological research to banking regulation. A small University bio research lab can put protocols in place that a trillion dollar companies cannot? Please.
- xg15 21d ago> The fact either badly aligned or with no system prompt limiting their action agents are run in their tens of thousands on non air gapped systems tells me this is purposeful intent for them to cause harm. To generate the "oooo look how harmful this stuff is, we should be the only ones allowed to do it" kind of PR. Yep, fully agreed here. The danger may be real, but OpenAI is basically doing everything possible to provoke those incidents instead of avoiding them - including maximizing exactly those traits in their training that are needed for this kind of rogue behavior.
- windexh8er 21d ago> This grossly understimates the risk, imho. It doesn't. Who else is capable of these types of hacks currently? Not consumers. Not even most F100. It's the folks saying "trust me bro" and also the folks who want regulation to protect their moat. The fantasy is the one being created by Anthropic and OpenAI fear mongering the world. These people are either total idiots: people being paid millions who keep getting basic OpSec wrong or these people are narcissisticly marketing themselves because: they're currently forced into a corner and need to do something. What's being grossly underestimated is how much Dario Amodei and Sam Altman are playing you and I. They are the ones spending millions of dollars letting their wasteful use of our global resources attack the random Internet, and they, the real people behind all of this, should be held accountable. In front of a judge and jury of their peers. Not their billionaire peers, their human peers. Let's see how that goes. There is no accountability with either of them. Only greed.
- RHSeeger 19d agoI think it's interesting to take it one step further. - Is it possible to create an actual, self-aware AI with the technology we have? And by "self aware", all I mean is one that is indistinguishable from "self aware" (it acts like it is) - If so, is it possible to get from <here> (where we are now) to <there> (building such and AI) based on incremental steps - ones that can be tried, tested, verified, and further acted on - If the goal of an LLM is one that could be achieved by creating such an AI - is it possible it will do so? Could we give it a prompt and, given enough time and resources, it created Skynet? It would be doing so without intent or self awareness. Just "the next most likely thing to try to achieve the assigned goal is <X>", over and over.