4 ms·
The AGI safety stuff is weird and silly. AFAICT that community is mostly filled with philosophy types anthropomorphizing ML models that they don't understand. I
by throwawaygh 4y ago
The AGI safety stuff is weird and silly. AFAICT that community is mostly filled with philosophy types anthropomorphizing ML models that they don't understand. It's a waste of time and money. But then, I feel the same way about like 80% of what NSF CISE funds and there's plenty of toxic types in academia. At least the AGI safety stuff is all private money. My whole reaction to that community is mostly "weird and not worth my time".
Timmit et al.'s reaction to the AGI/existential safety stuff is also weird. I don't even disagree, but I'm not sure why this topic in particular is such a lightening rod.
- petters 4y ago"anthropomorphizing ML models" seems like the complete opposite of what AI safety people do. They often point out that just because you are intelligent, you don't have to seem human. I think they call it the "orthogonality thesis".
- heavyset_go 4y agoAnthropomorphism isn't just about directly comparing things to people, it's also about assigning qualities that are related to people, like motivations, goals, opinions, concepts like good and bad, etc, even if the specific qualities would be alien to most people. In that sense, the AI safety crowd has been anthropomorphizing hypothetical AI for close to 20 years now.
- brian_cloutier 4y agoDo you have an example of someone in the community making this mistake? My understanding is that if we knew that all future AGIs would have human-like motivations, goals, opinions, and concepts of good and bad then that crowd would be much less concerned. Smarter humans are not the concern. Intelligences which are _not_ human-like are the explicit concern.
- heavyset_go 4y agoRoko's basilisk comes to mind.
- brian_cloutier 4y agoI wonder if you've had an in-person conversation with someone who would identify as being a part of the AGI safety community. I've talked with several people and their arguments are a lot more sophisticated than an intuition that ML models would have human motivations. This post isn't written by somebody inside the community but he presumably has access to them and has had conversations with them which shape his beliefs: https://astralcodexten.substack.com/p/deceptively-aligned-mesa-optimizers https://astralcodexten.substack.com/p/deceptively-aligned-me... I wonder if you also believe that post is just a wasteful collection of philosophy anthropomorphizing misunderstood ML models.
- SpicyLemonZest 4y agoThe core thing I (and I suspect the original commenter) struggle to get past is the invetiable twist that in this post happens halfway through this post's part II. > If it’s a very smart mesa-optimizer, it might think “If I throw the strawberry at the streetlight, I will be caught and trained to have different goals." It seems to me that this is a category error, like having the very smart mesa-optimizer start thinking about how it can find other models to marry. Why would gradient descent produce this very specific concept of goals-based identity? It's not even a human universal - many people don't have a particularly strong attachment to their current set of goals and hope for God or Buddha to help them get different ones.
- brian_cloutier 4y agoI don't think I understand your objection because it doesn't seem like a category error to talk about optimizers having goals. I think you would agree that thermostats have goals? They try to minimize the error between the desired and the actual temperature. And you would also agree that gradient descent has a goal? It tweaks parameters in the search for models which minimize error in the training set. The system performing that gradient descent was designed by humans and exhibits goal-like behavior. But you think it's a step too far to believe that gradient descent could create a model which also exhibits goal-like behavior? What is the difference in category that you see between those two steps? I agree that humans are not goal-directed in the same way the community is worried that AGI might be. This makes it surprising that seeing AGI as potentially goal-directed is seen as anthropomorphism, humans often question their goals in exactly the way there is concern that AGI will not!