5 ms·
I agree with you that "AI safety" (let's call it bickering) and "alignment" should be separate. But I can't stomach the thought experiments. First of all, it ta
by srslack 3y ago
I agree with you that "AI safety" (let's call it bickering) and "alignment" should be separate. But I can't stomach the thought experiments. First of all, it takes a human being to guide these models, to host (or pay for the hosting) and instantiate them. They're not autonomous. They won't be autonomous. The human being behind them is responsible.
As far as the idea of "hacking some funny Internet money, using it to mail-order some synthesized proteins from a few biotech labs, delivered to a poor schmuck who it'll pay for mixing together the contents of the random vials that came in the mail... bootstrapping a multi-step process that ends up with generic nanotech under control of the AI.":
Language models, let's use GPT-4, can't even use a web browser without tripping over itself. My web browser setup, which I've modified to use the chrome visual assistance over the debug bridge now, if you so much as increase the pixels of the viewport by 100 or so, the model is utterly perplexed because it's lost its context. Arguably, that's an argument from context, which is slowly being made irrelevant with even local LLMs (https://www.mosaicml.com/blog/mpt-7b https://www.mosaicml.com/blog/mpt-7b). It has no understanding, it'll use an "example@email.com" to try and login to websites, because it believes that this is its email address. It has no understanding that it needs to go register for email. Prompting it with some email access and telling it about its email address just papers over the fact that the model has no real understanding across general tasks. There may be some nuggets of understanding in there that it has gleaned for specific task from the corpus, but AGI is a laughable concern. These are trained to minimize loss on a dataset and produce plausible outputs. It's the Chinese room, for real.
It still remains that these are just text predictions, and you need a human to guide them towards that. There's not going to be autonomous machiavellian rogue AIs running amok, let alone language models. There's always a human being behind that.
As far as multi-modal models and such, I'm not sure, but I do know for sure that these language models don't have general understanding, as much as Microsoft and OpenAI and such would like them to. The real harm will be deploying these to users when they can't solve the prompt injection problem. The prompt injection thread here a few days ago was filled with a sad state of "engineers", probably those who've deployed this crap in their applications, just outright ignoring the problem or just saying it can be solved with "delimiters".
AI "safety" companies springing up who can't even stop the LLM from divulging a password it was supposed to guard. I broke the last level in that game with like six characters and a question mark. That's the real harm. That, and the use of machine learning in the real world for surveillance and prosecution and other harms. Not science fiction stories.
- kalkin 3y ago"First of all, it takes a human being to guide these models, to host (or pay for the hosting) and instantiate them" And this will always be true? You repeat this claim several times in slightly varied phrasing without ever giving any reason to assume it will always hold, as far as I can see. But nobody is worried that current models will kill everyone. The worry is about future, more capable models.
- srslack 3y agoWho prompted the future LLM, and gave it access to a root shell and an A100 GPU, and allowed it to copy over some python script that runs in a loop and allowed it to download 2 terabytes of corpus and trained a new version of itself for weeks if not months to improve itself, just to carry out some strange machiavellian task of screwing around with humans? The human being did. The argument I'm making is that there's actual real harms occurring now, not some theoretical future "AI" with a setup that requires no input. No one wants to focus on that, and in fact it's better to hype up these science fiction stories, it's a better sell for the real tasks in the real world that are producing real harms right now.
- pjc50 3y ago> Who prompted the future LLM, and gave it access to a root shell and an A100 GPU, and allowed it to copy over some python script that runs in a loop and allowed it to download 2 terabytes of corpus and trained a new version of itself for weeks if not months to improve itself, just to carry out some strange machiavellian task of screwing around with humans? > The human being did. I generally agree with you and think the doomerists are overblown, but there's a capability argument here; if it is possible for an AI to augment the ability of humans to do Bad Things to new levels (not proven), and if such a thing becomes widely available to individuals, then it would seem likely that we get "Unabomber but he has an AI helping him maximise his harm capabilities". > it's a better sell for the real tasks in the real world that are producing real harms right now. Strongly agree.
- kalkin 3y ago> The human being did. I'm not sure whether you're making an argument about moral responsibility ultimately resting with humans - in which case I agree - or whether you're arguing that we'll be safe because nobody will do that with a model smart enough to be dangerous - in which case I'm extremely dubious. Plenty of people are already trying to make "agents" with GPT4 just for fun, and that's with a model that's not actively trying to manipulate them. > actual real harms occurring now Sure, but it's possible for there to be real harms now and also future potential harms of larger scope. Luckily many of the same potential policies - e.g. mandating public registration of large models, safety standards enforced by third-party audits, restrictions on allowed uses, etc - would plausibly be helpful for both. > science fiction stories There's no law of nature that says if something has appeared in a science fiction story, it can't appear in reality.