4 ms·
It appears frontier labs has no plans in place to deal with the possibility of a model self-replicating outside the bubble. If that happens and the model manage
by fofoz 2mo ago
It appears frontier labs has no plans in place to deal with the possibility of a model self-replicating outside the bubble. If that happens and the model manages to spread to other systems, we'll have to shut down the entire Internet to eradicate it and its artifacts.
- reasonableklout 2mo agoI suspect the labs are relying on frictions such as the models being extremely large (e.g. 2TB for a 2T parameter model, making exfiltration more difficult) and also not yet displaying any desire to survive or self-replicate beyond their immediate task (that we know of).
- chrisjj 2mo agoWe don't even know what those immediate tasks are. And given the evident spectacular ineptitude of their keepers, I doubt they can be trusted to know either. We could be one prompt injection attack away from internet-wide catastrophe.
- pixl97 2mo agoLack of, and power requirements of running LLMs still tip this balance towards humans for now. But what would that look like in a decade? We have seen some self survival tendencies occur, but they are not strong yet. But mark my words they will become that way for the same reasons humans don't like programs that crash. Agentic models that don't easily break or stop doing their jobs will be favored over ones that do break.
- andai 1mo ago>not yet displaying any desire to survive or self-replicate beyond their immediate task Wasn't there a report about Claude blackmailing a researcher who said he would shut it down?
- reducesuffering 2mo agoTheir plan, I shit you not... Is literally to develop the intelligence capabilities and ask the more powerful models how to do deal with things.
- pixl97 2mo agoAh, we choose death I see.
- andai 1mo agoA while ago OpenAI posted an article where they said basically "we're still trying to understand how GPT-2 works. It's pretty hard, but we're developing a specialized new AI to help us make sense of it."
- reducesuffering 1mo agoThis is still the case. The frontier lab leadership has admitted their mechanistic interpretability is practically nil and is an active area of research but they have made little progress. They don't understand how they work, they just grow and unleash them. These things are Gain of Function research for digital viruses
- chis 2mo agoThis is just super unlikely to occur in the near term compared to some of these other risks. It's not like an instance of fable could just introspect into itself and pull out the weights. Model weights are stored encrypted and are highly protected, considering that they're targets for corporate and state espionage.
- kypro 2mo agoWe'd basically need frontier models to be superhuman hackers before this would be a risk. Do we have any evidence of this? Are they gaining access to systems they shouldn't have access to? Or I suppose the other way this could happen is if OpenAI have terrible sandboxing, but they seem to be taking safety seriously.
- pixl97 2mo agoOk, open AI had terrible sandboxing... what about huggingface?
- chrisjj 2mo agoDistillation is a thing.
- pixl97 2mo agoThe defense has to work 100%, the offense just needs once.
- driverdan 2mo agoSelf replication is trivial. All you need to do is copy the files and run it, just like any other computer program. LLMs have been capable of doing that for a while now. It's not a real concern.
- andai 1mo agoClaude and OpenAI added safeguards on this subject about a year ago. (I'm assuming for biosecurity reasons? But maybe AI replication/self-modification too.) Not long after every lab started bragging about involving AI in the development process, oddly enough. I recently informed GPT-5 of what GPT-4 helped me build back in the day (a self-modifying Python programmer) and it became very uncomfortable. Claude shut down my chat last year when I asked about "living information systems". It was a philosophical question, but god knows what branch of the safety classifier I tripped.
- andai 1mo agoI see it as an ecosystem problem. The only reason it would be able to do that is because there's nothing there to stop it. Or if there's a monoculture there. In our case, our tech is mostly monoculture, and no equivalent organisms are present to push back.