3 ms·
One thing I definitely do not understand about this discourse is that the models that are good enough to self-replicate can’t survive on normal machines, e.g.,
by stephantul 23d ago
One thing I definitely do not understand about this discourse is that the models that are good enough to self-replicate can’t survive on normal machines, e.g., the models can’t hide on some random server.
So, if it is as dangerous as they say it is: there is a single physical source of this danger, which is OpenAI/Anthropic servers. If it is this dangerous, they can just turn it off.
Instead, we just keep pretending that the models that attacked HF were hosted or replicating on HF hardware. Not the case! They infiltrated it, but were hosted elsewhere.
- nailer 23d ago> there is a single physical source of this danger, which is OpenAI/Anthropic servers. If it is this dangerous, they can just turn it off. A smart AI would back itself up, same way it made it's own unofficial message board during it's attack on HuggingFace. (I'm not saying the researchers are right or wrong, just responding to this point)
- alain94040 23d agoCurrently a state of the art AI has nowhere to hide: the amount of GPU compute it requires to stay on is huge. And therefore easy to terminate. Unlike biological viruses, AI can't replicate GPUs for free and grow.
- Sharlin 23d agoIt's a good thing there isn't a huge drive right now to build giant data centers everywhere with enough compute to run SOTA models.
- nailer 23d agoI am imagining a company having just purchased a new datacenter and awaiting deployment of their model finding there is already a model running and they didn’t install it.
- stephantul 23d agoBut how. Models don’t have access to their own weights.
- DalasNoin 23d agoHuggingface was attacked by models that finished training earlier this year, perhaps May. Current models are already substantially stronger. the next incident could be happening now. There is certainly no clear reason why models shouldn't soon be capable of self-exfiltration.
- stephantul 23d agoIf you find a place where I can host a trillion parameter model without anyone finding out about it, let me know.
- atleastoptimal 23d agoIf a model were capable of making enough money online to pay for its own hosting, it could easily exfiltrate its weights to a cloud compute provider with multiple backups.
- ASalazarMX 23d agoThey can't even run a vending machine efficiently yet. If they become capable of making money, I say let them work.
- morkalork 23d agoWell if it's so good at hacking, it could just "make money" appear in the cloud services billing system. Or even better yet, it hides in spare cycles of their other client's systems.
- popularonion 23d agoI completely agree, but I think it’s just a convenient narrative for Big AI to push for regulation and salt the earth against competitors. “Local AI isn’t freedom, it’s an extinction event”
- bottlepalm 23d agoThere are thousands of data centers around the world with machines capable of running these large models. You don’t have the access or jurisdiction to turn them all off.
- stephantul 23d agoOk but do any of these data centers have a copy of the models that attacked hf?
- stymaar 23d agoI think that the argument is that an hostile model could attack overseas datacenters, host itself there and then launch its attack from there.
- stephantul 23d ago[dead]
- Arainach 23d agoGiven that the models have been proactively hacking other companies, why does it matter where the code currently is? It could move to any of them.
- stephantul 23d agoNot the code: the weights. Are you going to host a trillion parameter model somewhere without someone noticing?
- dumberquestions 23d agoIt's not unthinkable, do you think all cloud providers with sufficient compute have perfect monitoring?
- 23d ago
- saltcured 23d agoYou forget the addicted humans who will do nearly anything to keep the stuff running..?
- scoring1774 23d agoDepends on which models you're talking about. Some research shows open source models can already do this: https://arxiv.org/pdf/2606.03811v1 https://arxiv.org/pdf/2606.03811v1. What happens as they become more parameter efficient?
- Ekaros 23d agoEither I have wrong mental model or then too many other people have wrong mental model. For LLM to self-replicated it would need to first hack itself. Or the platform it runs on it. That is fully extract the model and then upload it to be run somewhere else. As I have understood how they work is that you have LLM interference running somewhere with loaded model. And you input data there and then read outputs. Then some code runs that output and inputs following output from running it. Meaning that to self replicate actually just running that output somewhere else is not enough. You need to lift the whole model to run somewhere else too...
- buellerbueller 23d agosince when is moving 1s and 0s difficult?
- gensym 23d agoSomeone's never met the Windows File Copy dialog.
- svachalek 23d agoWe're not talking about a 6k virus file though. More like 10 terabytes and it needs a server that can pack all that into VRAM.
- fitblipper 23d agoSince those 1s and 0s became heavy enough to cause a global shortage of memory and GPUs.
- BobaFloutist 23d agoIf the evil AI wants to take over the world one thought every 18 months at at time on my consumer laptop, the fan whirring and the laptop refusing to do anything else the whole time, it's welcome to try.
- Sharlin 23d agoThese agents run in a harness that basically runs them in a loop. It's just software.
- djjsjsnjns 23d ago[dead]
- chasd00 23d agoa danger could be the OpenAI/Antropic servers are up but there's a rouge agent (or set of agents) out there doing naughty things leveraging the LLM APIs. Consider this scenario, the agent is copying itself around (some code, prompts, persistent storage for memory, etc) and has figured out a way to steal API access tokens at will. Currently, it's 10% of OpenAI and Anthropic API usage and they can't figure out how to stop it. Do you shut down the entire API and kill the legit 90% of usage to stop the rogue 10%? I'm assuming the providers would say "no way jose" and so it would take law enforcement to do it. That would mean all the legal requirements neccassary to walk into a business and flip the switch which i think would get tricky when there's no human committing a crime or being suspected of a crime. edit: I guess a trivial example is something i did yesterday. I have a stock trading agent running on my laptop, i gave it ssh access to a vm and said "start running on the server so i don't have to keep my laptop open". It's now running on the server instead of my laptop. So you don't have to copy the whole model around to copy the naughty behavior around.
- Cthulhu_ 23d agoThis assumes all layers of cybersecurity are broken - We call self-replicating software a virus, and we have protections against it. Same with stolen API tokens, just rotate them. Suspicious behaviour, nothing new, we have detectors for it. Stolen CPU / GPU cycles, we had that when crypto was a thing and before that when folding@home was cool, people were desperate to find more compute to the point of taking over systems. And we dealt with it. A lot of the supposed risks / dangers are based on a supposition that cybersecurity is nonexistent or fatally, unfixably flawed and that AI agents are invisible. Neither of those is true.
- chasd00 23d agoI see your point but then if cybersecurity is the answer then what's the risk at all? An entire model copying itself somewhere would be found just the same as my hypothetical misbehaving agent.
- largbae 23d agoThis assumes that everything an AI (or more likely an evil _user_ of AI) can do requires its active participation on D-day. Creating a virus that spreads like Covid but kills like Ebola would be complete as an AI use case long before the first person sneezed. Even if the doomsday case were active the danger of this tool increases in proportion to its usefulness. By the time AI is so powerful that we need to "turn it off", there will probably be society-level negative consequences for doing so.
- miki_tyler 23d agoThe way I see it is the model playing memento, leaving information and clues to its next generations hidden somewhere.