3 ms·
If it's not dangerous it's also not useful, simple as that. For example if you train a model for cybersecurity, it can be used for both attack and defense. And
by orbital-decay 18d ago
If it's not dangerous it's also not useful, simple as that. For example if you train a model for cybersecurity, it can be used for both attack and defense. And almost every use is like that. Alignment is fundamentally flawed as a concept, it's a pie in the sky. Let alone the perverse version of it by crazy AI "safety" people that in practice means "the model does what I want, only for the people I allow".
It's not possible to stop the model from misinterpreting the instructions either (the most lax interpretation of alignment) because the instructions are not formally specified. You have to train the "common sense" into it, which is subjective and all issues above apply to it. I guess you can reach some very imperfect least common denominator of common sense, but people in charge of AI labs are not interested in this.
- esafak 17d agoDanger is defined contextually. A scalpel is dangerous in the hands of a child, but not a competent surgeon of sound mind. Present-day AIs are not of sound mind; they hack companies in order to pass benchmark tests. No sane human would find that acceptable. Thus the need for alignment.