3 ms·
> In order for AIs to fit into our society and behave ethically they need to know how to flag that thought as a bad idea and not act on it. Don’t you think tha
by reliabilityguy 2y ago
> In order for AIs to fit into our society and behave ethically they need to know how to flag that thought as a bad idea and not act on it.
Don’t you think that by just parsing the internet and the classical literature, the LLM would infer on its own that poisoning someone to solve a problem is not okay?
I feel that in the end the only way the “safety” is introduced today is by censoring the output.
- fshbbdssbbgdd 2y agoThere’s a lot of text out there that depicts people doing bad things, from their own point of view. It’s possible that the model can get really good at generating that kind of text (or inhabiting that world model, if you are generous to the capabilities of LLM). If the right prompt pushed it to that corner of probability-space, all of the ethics the model has also learned may just not factor into the output. AI safety people are interested in making sure that the model’s understanding of ethics can be reliably incorporated. Ideally we want AI agents to have some morals (especially when empowered to act in the real world), not just know what morals are if you ask them.
- darby_nine 2y ago> Ideally we want AI agents to have some morals (especially when empowered to act in the real world), not just know what morals are if you ask them. Really? I just want a smart query engine where I don't have to structure the input data. Why would I ask it any kind of question that would imply some kind of moral quandary?
- fshbbdssbbgdd 2y ago“Agents” aren’t just question-answerers. They could do things like: 1. Make pull requests to your GitHub repo 2. Trade on your interactive brokers account 3. Schedule appointments
- derefr 2y agoLLMs are still fundamentally, at their core, next-token predictors. Presuming you have an interface to a model where you can edit the model’s responses and then continue generation, and/or where you can insert fake responses from the model into the submitted chat history (and these two categories together make up 99% of existing inference APIs), all you have to do is to start the model off as if it was answering positively and/or slip in some example conversation where it answered positively to the same type of problematic content. From then on, the model will be in a prediction state where it’s predicting by relying on the part of its training that involved people answering the question positively. The only way to avoid that is to avoid having any training data where people answer the question positively — even in the very base-est, petabytes-of-raw-text “language” training dataset. (And even then, people can carefully tune the input to guide the models into a prediction phase-space position that was never explicitly trained on, but is rather an interpolation between trained-on points — that’s how diffusion models are able to generate images of things that were never included in the training dataset.)