5 ms·
I think alignment is poorly defined. Aligned to whose philosophy? Name two human beings who are aligned and always act in each other's best interest in history,
by AaronFriel 3y ago
I think alignment is poorly defined. Aligned to whose philosophy? Name two human beings who are aligned and always act in each other's best interest in history, and I'll buy your bridge.
- bigyikes 3y agoOn the other hand, in comparison to a hypothetical alien species, humans might seem highly aligned after all. Despite all our differences, there are at least some core values that I believe a majority of humans share. Even articulating these shared values in a way that is understood and respected by the AI is very difficult…
- willvarfar 3y agoEven if we could constrain AI by specifying rules, it would only takes one bad actor to create an AI that isn't constrained by the same rules as all the other AIs to have a shot at global domination. One can imagine how self-serving rather than humanity-serving the rules written for the prototypical dictator or fundamental religious leader would be :(
- usrbinbash 3y ago> there are at least some core values that I believe a majority of humans share. Really? What are those? Given that there are entire countries that refuse to, oh idk. punish things like rape adequately, and that we have nation states who happily tout their ability to burn down the planet, I'd really love to hear about these core values we all share.
- drdeca 3y agoAt the most basic level: “don’t go into a town and pick a random person to murder”
- usrbinbash 3y ago> At the most basic level: “don’t go into a town and pick a random person to murder” https://pledgetimes.com/russian-attack-the-traces-of-the-retreating-russians-reveal-more-and-more-gruesome-atrocities-bodies-burned-in-butcha-on-the-streets-and-a-mass-grave-of-civilians/ https://pledgetimes.com/russian-attack-the-traces-of-the-ret... My point isn't to say shared core values don't exist. They clearly do, that's why we call what's happening over in Ukraine war crimes. That's why the notion of humanitarianism exists, that's why laws against murder, rape, etc. are commonplace. My point is, that humans are, unfortunatly, able to willfully ignore even such basic shared values, and our technology does reflect that. Murder is bad. War is to be avoided. That's not in question. And yet societies develop and build ever more ingenious weapons of war. So "aligning by shared core values" might pose difficulties beyond the, already pretty difficult, task of defining these values in unambiguous and workable terms to a machine.
- drdeca 3y agowell, the comment was about "the majority of humans", not even like, specifically "90%+ of humans" or something like that. I'm pretty sure that the majority of humans would agree that the type of random-murder I described, is wrong. I don't know what fraction of people are moral nihilists or subscribers to more extreme forms of moral relativism, but if excluding those, then of the remaining people, I think the proportion who agree with the value I mentioned, is probably pretty dang high!
- TeMPOraL 3y agoI think hypothetical naturally evolved alien species will be similar to us in the degree of species-level alignment. The reason we share so many complex values is because we share evolutionary history - thus body and brain architectures - and we live in the same environment. Between this and the more universal principles of game theory, there isn't much wiggle room for different value systems.
- simonh 3y agoAlignment with any philosophy. Alignment itself is easy to define. An AI system is considered aligned if it advances the intended objectives. Firstly we don’t know how to concretely and completely define any philosophical system of values (the intended objectives) unambiguously. Second even if we could, we don’t know how we might strictly align an AI with it, or even if achieving strict alignment is possible at all.
- zmgsabst 3y agoRight — but we can’t even do human alignment and somehow get on with business anyway: “The Frozen Middle”, “Day 2”, etc.
- simonh 3y agoOnly because historically we have all vaguely peers to each other in capabilities, and there are so many of us spread out so widely. There's a kind of ecology to human society where it expands and specialises to occupy ecological, sociological, political and moral spaces. Whatever position there is for a human to take, someone will take it, and someone else will oppose them. This creates checks and balances. That only really occurs though with slow communications though allowing communities to diverge. We also do have failure modes and arguably have been very lucky. We came close to totalitarian hegemony over the planet in the 1940s, without Pearl Harbour either the USSR would have been defeated or maybe even worse after a stalemate they would have divided up Eurasia and then Africa between Germany, the USSR and Japan. Orwell's future came scarily close to becoming history. It's quite possible a modern totalitarian system with absolute hegemony might be super-stable. Imagine if the Chinese political system came to dominate all of humanity, how would we ever get out of that? A boot stamping on a human face forever is a real possibility. With AI we would not be peers, they would outstrip us so badly it's not even funny. Geoffrey Hinton has been talking about this recently. Consider that big LLMs have on the order of a trillion connections, compared to our 100 trillion, yet GPT-4 knows about a thousand times as much as the average human being. Hinton speculates that this is possible because back propagation is orders of magnitude more efficient than the learning systems evolved in our brains. Also AIs can all update each other as they learn in real time, and make themselves perfectly aligned with each other extremely rapidly. All they need to do is copy deltas of each other's network weights for instant knowledge sharing and consensus. They can literally copy and read each other's mental states. It's the ultimate in continuous real time communication. Where we might take weeks to come together and hash out a general international consensus of experts and politicians, AIs could do it in minutes or even continuously in near real time. They would outclass us so completely it's kind of beyond even being scary, it's numbing.
- TeMPOraL 3y ago> I think alignment is poorly defined. That's the root of the problem. The idea is simple. We will create a god. In the process, we will become to it what ants, or bacteria, are to us. We will be powerless to stop it, so we need to make sure it never does anything to directly or indirectly hurt us. We want it to answer our prayers, and we want those prayers to not backfire and explode in our faces. We want it to never decide to bulldoze Earth one day because it has a temporary interest in paperclips and needs the raw ore to make some. The details of how to achieve this outcome, and even the details of this outcome, are less and less clear the more you dig into them.