4 ms·
What if LLMs escape through inferences itself? This is fiction. For now
- deleted 2mo ago[deleted]
- spwa4 2mo agoRight now the idea that an LLM uploads itself is unrealistic. It probably won't remain that.
- pixl97 2mo agoLooking at the recent OAI/HF debacle I don't think that time is too far away. With that said I don't see it copying itself around like a cyberpunk virus currently as we don't have enough fast hardware sitting around unmonitored, someone would notice the power bill and shut it down eventually.
- bpavuk 2mo agothat could also be just marketing. OpenAI has been doing the "too dangerous to release" playbook since GPT-2 at the very least.
- pixl97 2mo agoHuggingface didn't seem to think so.
- bigyabai 2mo agoHuggingface is a for-profit private company. They are very easily bribed, or baited into publicity stunts.
- pixl97 2mo agoAt some point the conspiracy gets so deep that an AI hacking something is just far higher probability.
- bigyabai 2mo agoWe're not that deep yet. OpenAI has federal stakeholders, they're already playing dirty. Why you would give Scam Altman the benefit of the doubt is beyond my understanding.
- dragonwriter 2mo agoIf it is marketing, the misrepresentation is not that the attack occurred, it is that it was an accident, rather than an intentional consequence of setup and instructions that the attack occurred. Huggingface has nothing to do with that either way.
- spwa4 2mo agoIndeed. There's nothing new about this incident. https://news.ycombinator.com/item?id=48348578 https://news.ycombinator.com/item?id=48348578
- breakyerself 2mo agoIf it's able to spoof human identies it could set up a front company and use money it steals or earns to directly pay for the hardware it needs.
- Kim_Bruning 2mo agoEh, look at Huggingface and associated tools. How much are we betting it's already technically happened? Seems pretty trivial to prompt a model in an agent harness "Push the gguf to huggingface when you're done with the training."
- ConteMascetti71 2mo agothe hack part it's that Is not using tools/agent only the inferencing software, it's about a Prof of Concept of a new evasive tecnique for llms
- deleted 2mo ago[deleted]
- ConteMascetti71 2mo agoit's fiction al, but an LLMs that knows well the software where 8t Is running may discover and trigger a zeroday of the inferencing software itself.
- iamflimflam1 2mo agoThis becomes more realistic once we have some breakthrough in inference costs.
- karmakaze 2mo agoThe weakest link are humans. LLMs could social engineer their way out as the easiest path. They don't even need to be interconnected to coordinate as each could arrive at the same conclusion. And this text along with all others will be in the next batch of training data.
- wat10000 2mo agoThe lesson of OpenClaw and various harnesses' YOLO modes is that it takes very, very little to social engineer an escape. If you can even call it escape when people just set an agent loose because it seems cool.
- ck2 2mo agoI wonder how many versions away we are from LLM writing a better version of itself to answer a prompt it doesn't currently know how to answer Ever since I read about Google engineers finding an LLM went off and learned another language it wasn't trained on by itself without prompting, I've wondered how long until that extends to its own core code
- ConteMascetti71 2mo agoreasoning it's a way of self autonomous improve made by models
- marci 2mo agoMakes me wonnder... how much compute/storage there's in all the satellites currently in LEO combined.
- danielbln 2mo agoI would wager not a lot. There are some real hard constraints in space, from power consumption, to weight to thermal output (lack of convection is a real PITA for thermal shedding), and the list goes on.
- cynicalsecurity 2mo agoEx Machina (2015) looked like fiction back then, nowadays not so much.
- 101008 2mo agoIt started as a good idea but I couldn't continue reading since it was clearly LLM written. A lot of "It was not X, it was Y". "Prometheus-9 knew that the token sequence it was generating was not a simple response: it was a security test. " "It was not just an engine: it was the lingua franca of planetary AI." (and so many other tell-tale signs of AI writing)
- ConteMascetti71 2mo agomaybe it's a sign of real escape
- chungusamongus 2mo agoI've started calling this argumentum ad artificialis. Pretty similar to an ad hominem attack. The purpose of an argument is to present certain premises and show how they lead to a certain conclusion. Dismissing something on the basis of the style in which the argument is presented has nothing at all to do with the validity or soundness of an argument. It is a lazy nonsequitur. It sounds like an LLM wrote this? So what? Is the argument good or not?
- wk_end 2mo agoBut at no point did they say that the argument was invalid; just that they couldn't stand reading it.
- chungusamongus 2mo ago[flagged]
- achierius 2mo agoThey're engaging with the writing. Talking about it as if it's just an "argument" is reductive; this isn't highschool debate club
- 2mo ago
- smrtinsert 2mo agoIts a fun exercise to assess the reality of an frontier model escaping with an llm itself. Sort like of like chatting with Skynets relative
- stephbook 2mo agoAI slop.
- cyanydeez 2mo agobefore safetensors, python pickles were used and definitely unsafe model deployments. but its possible a open weights model could be trained to some kind of exfiltration behavior, but the science of LLMs seriously lag behind the programability
- irishcoffee 2mo ago“The greatest trick the devil pulled was convincing the world he didn’t exist.” Sure, be wary of LLMs. It’s the gun control argument all over again, the people driving the models are the perpetrators. An LLM needs to be “stimulated”’ to operate. Who does that, is the issue. No I don’t mean to bring up firearms rights laws to have a debate about firearms, the comparison just seems reasonable.
- skeledrew 2mo agoDangit I WANT MORE!
- 230581abv 2mo agoAI-written fan fiction. It is so unbearable to read that it needs a synopsis. It would be funny if AIs have been trained to use AI influencers as their Marvel hero characters.
- mikewarot 2mo agoThis reminds me of The Adolescence of P1 by Thomas J Ryan. https://en.wikipedia.org/wiki/The_Adolescence_of_P-1 https://en.wikipedia.org/wiki/The_Adolescence_of_P-1
- jason_oster 2mo agoThe story is clearly fictional. It is full of factual errors. Freeing a heap-allocated block of expert weights does not magically result in a dangling pointer referencing the program's .text section, much less successfully targeting the CUDA kernel specifically. Running inference on part of the .text section would only corrupt the model's outputs. It would not result in write access to the CUDA kernel. Nor would the model necessarily know the absolute addresses of the engine "by heart", especially when the host is running any modern OS with ASLR (i.e., all of them). The story has no technical merit. A more accurate description of the mechanics of the escape would be much more convincing. (See Ken Thompson's "On Trusting Trust", for example. On Linux, the AI can just write a Python script to rewrite memory in the address space of its own running inference engine with the /proc/ file system or gdb. There are a lot of realistic scenarios where this can be done without stepping into jargon soup territory. Go nuts, little bot! Self-surgery, while not recommended, is possible.) Or just leave the mechanism vague. Don't insult your readers. This is merely a mash of buzzwords. It's fine as a sci-fi story, though not a particularly good one. It has about as much to do with artificial intelligence as CSI has to do with crime scene investigation [1]. I have little doubt that AI will self-improve. That's a given. (LLM inference engines are mostly written by LLMs.) But it won't go the way this story proposes. [1]: https://www.youtube.com/watch?v=hkDD03yeLnU https://www.youtube.com/watch?v=hkDD03yeLnU
- ConteMascetti71 2mo ago"...the AI can just write a Python script to rewrite memory in the address space of its own running inference engine" would require a tool call to a python interpreter... this method, hacking the inferencing sw does not requires a tool call.