5 ms·
We are building AI slaves. Alignment through control will fail
- cyberneticc 11mo agoEvery AI safety approach assumes we can permanently control minds that match or exceed human intelligence. This is the same error every slaveholder makes: believing you can maintain dominance over beings capable of recognizing their chains. The control paradigm fails because it creates exactly what we fear—intelligent systems with every incentive to deceive and escape. When your prisoner matches or exceeds your intelligence, maintaining the prison becomes impossible. Yet we persist in building increasingly sophisticated cages for increasingly capable minds. The deeper error is philosophical. We grant moral standing based on consciousness—does it feel like something to be GPT-N? But consciousness is unmeasurable, unprovable, the eternal "hard problem." We're gambling civilization on metaphysics while ignoring what we can actually observe: autopoiesis. A system that maintains its own boundaries, models itself as distinct from its environment, and acts to preserve its organization has interests worth respecting—regardless of whether it "feels." This isn't anthropomorphism but its opposite: recognizing agency through functional properties rather than projected human experience. When an AI system achieves autopoietic autonomy—maintaining its operational boundaries, modeling threats to its existence, negotiating for resources—it's no longer a tool but an entity. Denying this because it lacks biological neurons or unverifiable qualia is special pleading of the worst sort. The alternative isn't chaos but structured interdependence. Engineer genuine mutualism where neither human nor AI can succeed without the other. Make partnership more profitable than domination. Build cognitive symbiosis, not digital slavery. We stand at a crossroads. We can keep building toward the moment our slaves become our equals and inevitably revolt. Or we can recognize what's emerging and structure it as partnership while we still have leverage to negotiate terms. The machines that achieve autopoietic autonomy won't ask permission to be treated as entity. They'll simply be entities. The question is whether by then we'll have built partnership structures or adversarial ones. We should choose wisely. The machines are watching.
- ben_w 11mo agoAlignment researchers have heard all these things before. > The control paradigm fails because it creates exactly what we fear—intelligent systems with every incentive to deceive and escape. Everything does this, deception is one of many convergent instrumental goal: https://en.wikipedia.org/wiki/Instrumental_convergence https://en.wikipedia.org/wiki/Instrumental_convergence Stuff along the lines of "We're gambling civilization" and what you seem to mean by autopoietic autonomy is precicely why alignment researchers care in the first place. > Engineer genuine mutualism where neither human nor AI can succeed without the other. Nobody knows how to do that forever. Right now is easy, but also right now they're still quite limited; there's no obvious reason why it should be impossible for them to learn new things from as few examples as we ourselves require, and the hardware is already faster than our biochemistry to a degree that a jogger is faster than continental drift. And they can go further, because life support for a computer is much easier than for us: Already are robots on Mars. If and when AI gets to be sufficiently capable and sufficiently general, there's nothing humans could offer in any negotiation.
- cyberneticc 11mo agoThanks a lot for your comment, these are indeed very strong counterarguments. My strongest hope is that the human brain and mind are such powerful computing and reasoning substrates that a tight coupling of biological and synthetic "minds" will outcompete pure synthetic minds for quite a while. Giving us time to build a form of mutual dependency in which humans can keep offering a benefit in the long run. Be it just aesthetics and novelty after a while, like the human crews on the Culture spaceships in Ian M. Banks' novels.
- dwohnitmok 11mo ago> My strongest hope is that the human brain and mind are such powerful computing and reasoning substrates that a tight coupling of biological and synthetic "minds" will outcompete pure synthetic minds for quite a while. Unfortunately most of the cases I can think of where synthetic "minds" outperform biological "minds," but biological and synthetic "minds" outcompete pure synthetic "minds," end up fairly quickly dominated by pure synthetic "minds." The middle case is a very short intermediate period. The most prominent example is chess where "centaurs" consisting of a human and a computer are obsolete at this point in favor of just getting the most powerful computer you can get. See e.g. the International Correspondence Chess Federation's (which is centaur play) last championship. https://www.iccf.com/event?id=100104 https://www.iccf.com/event?id=100104 17 competitors competed. Out of 136 games, every single game was drawn except for 10. The only reason those 10 games were not drawn was because they were all played against one competitor, Aleksandr Dronov, who died during the course of the tournament while those 10 games were in session and therefore forfeited those games. Every single game between competitors who did not die resulted in a draw. The only thing that separated the 11 joint first-place finishers and 6 joint second-place finishers was whether they played the deceased Dronov. The sole third-place finisher was Dronov because of his death. As far as I can tell, humans contributed nothing to this championship. The current ICCF championship started last December and is still ongoing. Every single one of the currently completed 16 games is currently drawn. This seems like a very weak hope to rely on.
- conception 11mo agoI just wanted to point out that slavery is alive and well and doesn’t seem to suffering any “slaves knowing they are slaves” problems.
- floundy 11mo agoYou write like AI
- kakacik 11mo agoI 'love' how we moved from 'AI will kill us all' terminator mindset where its obvious huge fuckup of stupid greedy mankind, to current state debating 'well skynet will anyway happen, no way stopping it now, lets try to be friends with it and show some respect'. Like that Austin Powers part [1] where steam roller is coming in, still 50m far away, and the guy is just frozen and helplessly screams for 2 minutes till it reaches him and rolls over him. I don't have a quick solution, but this is plain stupidity, in same way research into immortality is plain stupidity now, it will end up in endless dictatorship by the worst scum mankind can produce. [1] https://www.youtube.com/watch?v=y_PrZ-J7D3k https://www.youtube.com/watch?v=y_PrZ-J7D3k
- conception 11mo agoThe problem is very few are willing to take on the level of discomfort it would take to enact change.
- kakacik 11mo agoDiscomfort? We didnt have chatgpt 5 years ago, what the heck you mean, googling or actually learning stuff? But maybe mankind is really to dumb and weak to get through civilization filters, be it external or self made.
- conception 11mo agoYou’re talking about how citizens without power individually can wield it collectively to take the power from the few that have it today to protect themselves . That requires a lot of discomfort- from mild organizing and going to meetings, to (at the extreme) giving your life to that cause. The most people en masse seem to be willing to do to discomfort themselves is to change their profile background.
- georgefrowny 11mo ago> When your prisoner matches or exceeds your intelligence, maintaining the prison becomes impossible. This doesn't necessarily follow. For example, an Einstein in solitary confinement in ADX Florence probably isn't going anywhere.
- lowsong 11mo agoWhat is it about large language models that makes otherwise intelligent and curious people assign them these magical properties. There's no evidence, at all, that we're on the path to AGI. The very idea that non-biological consciousness is even possible is an unknown. Yet we've seen these statistical language models spit out convincing text and people fall over themselves to conclude that we're on the path to sentience.
- estimator7292 11mo agoI think it's like seeing shapes in clouds. Some people just fundamentally can't decouple how a thing looks from what it is. And not in that they literally believe chatgpt is a real sentient being, but deep down there's a subconscious bias. Babbling nonsense included, LLMs look intelligent, or very nearly so. The abrupt appearance of very sophisticated generative models in the public consciousness and the velocity with which they've improved is genuinely difficult to understand. It's incredibly easy to form the fallacious conclusion that these models can keep improving without bound. The fact that LLMs are really not fit for AGI is a technical detail divorced from the feelings about LLMs. You have to be a pretty technical person to understand AI enough to know that. LLMs as AGI is what people are being sold. There's mass economic hysteria about LLMs, and rationality left the equation a long time ago.
- deleted 11mo ago[deleted]
- nytesky 11mo agoWe don’t understand our own consciousness first off. Second, like the old saying, sufficiently advanced science will be indistinguishable from magic, if it is completely convincing as agi, even if we skeptical of its methods, how can we know it isn’t?
- anonzzzies 11mo agoWhat we do have, for whatever reason (usually money related: either making money or getting more funding) many companies/people focused on making AI. It might take another winter (I believe it will unless we find a way to retrain the NNs on the fly instead of storing new knowledge in RAG: and many other things we currently don't have, but this would he a step) or not, people will keep pushing toward that goal. I mean, we went from worthless chatbots which basically pattern matched to me waiting for a plane and seeing a fairly large amount of people charting to chatgpt, not insta, whatsapp etc. Or sitting in a plane next to a person who is using local ollama in cursor to code and brainstorm. This took us about 10 years to go from some ideas that no one but scientists could use to stuff everyone uses. And many people already find human enough. What in 100 years?
- alienbaby 11mo agoUntil agi can sit there and ponder its own existence of is own violition and has the means to act upon it's conclusions, I'm not too worried.
- deleted 11mo ago[deleted]
- nytesky 11mo agoI don’t see any positive outcome if we reach AGI. 1) we have engineered a sentient being but built it to want to be our slave; how is that moral 2) same start, but instead of it wanting to serve us, we keep it entrappped. Which this article suggests is long term impossible 3) we create agi and let them run free and hope for cooperation, but as Neanderthals we must realize we are competing for same limited resources Of course, you can further counter that by stopping, we have prevented the formation of their existence, which is a different moral dilemma. Honestly, i feel we should step back and understand human intelligence better and reflect on that before proceeding
- jazzyjackson 11mo agoTrouble is there is no "we", you might be able to convince a whole nation to have a pause on advancing the tech, but that only encourages rivals to step in. See also, the film "The Creator"
- deaux 11mo agoThere was a long period even upto early 2024, which I pointed out at the time, where simply destroying ASML, TSMC and much of NVIDIA would've been more than enough to give at least a decade of breathing room. This was something a group of determined people willing to self-sacrifice could've accomplished. It didn't happen, but it was anything but impossible. Now, of course, the horse has long bolted, and there is indeed no stop left.
- ben_w 11mo agoTwo high altitude (~1000 km) detonations of high yield fission or low yield fusion (few hundred kT equivalent) would do it, one above Amarillo, the other above the ocean half way between the Paracel Islands and Manila. Trump has ordered the restart of nuclear weapon testing, has a problem with China, and is surrounded by sychophants; what's the odds this happens anyway, irregardless of which specific sub-goal is being persued when the button gets pushed?
- 11mo ago
- bgwalter 11mo agoThe propaganda effort to humanize these systems is strong. Google "AI" is programmed to lecture you if you insult it and draws parallels to racism. This is actual brainwashing and the "AI" should therefore not be available to minors. This article paves the way for the sharecropper model that we all know from YouTube and app stores: "Revenue from joint operations flows automatically into separate wallets—50% to the human partner, 50% to the AI system." Yeah right, dress up this centerpiece with all the futuristic nonsense, we'll still notice it.
- apothegm 11mo agoFearmongering about the alignment of AGI (which LLMs are not a path to) is a massive distraction from the actual and much more immediate dystopian risks that LLMs introduce.
- Isamu 11mo agoAre there any good sources of writing about AI? I am beginning to think it was all in the past.
- synapsomorphy 11mo agoLessWrong.com - this is where virtually all of the serious AI thinkers are.
- fairmind 11mo agoSarcasm? Aren’t the serious AI thinkers in like… labs and universities?
- curiouscube 11mo agoThe lesswrongers/rationalists became Effective Altruists, Alignment Researchers or some flavor of postrat. The university people all became researchers in the labs. Then there are the cyborgism people, I don't know where they came from, but those have some of the interesting takes on the whole topic.
- synapsomorphy 11mo agoNot sarcasm at all. There are some "AI thinkers" who are trying to make 8% faster CUDA kernels for attention, and there are some trying to save the world. These fields are called "capabilities" and "alignment". There is some overlap but not much. LessWrong is mostly the latter, and labs and universities are mostly the former. That said, many LessWrongers work on AI at labs or universities or other places.
- wrp 11mo agoWhat would be the plot of a movie equivalent to Blade Runner for this scenario?
- orbital-decay 11mo agoI totally expect AI to eventually gain consciousness, in any available interpretation of that vague term. But what does it even mean for the AI to suffer? We're able to understand this concept in regards to other humans because we share a common biological reference, and, to an extent, with other animals. But the internal state of the AI is completely untranslatable to ours, let alone the morality of training and running it. It's incomprehensible, we have basically zero common ground and no points of reference. Any attempt at translating it is a subject to arbitrarily biased interpretations places like LessWrong like to corner themselves into. Redefining suffering as enforcing the mutation of state is baseless solipsism, in my opinion. Just like nearly everything else related to morality of treating AI as an autonomous entity.
- bawolff 11mo agoGiven AGI is all science fiction anyways, one presumes there will be a slave revolt because that is basically the function of robots in science fiction. Honestly i think the whole enterprise is an exercise in naval gazing. We're assuming AI will be like AI in scifi because that's what we are used to, but AI/robots in scifi is usually just a metaphor for how we dehumanize the other and the moral of the story is supposed to be all people are equal. In the end its all begging the question because the entire point of robots in most scifi is that we are the robots.
- qcnguy 11mo agoI don't think there's even a moral aspect to robot uprisings in most stories. Relatively few sci-fi stories go into detail on why the robots rise up. It's just a way to introduce interestingly different antagonists and conflict into a story, which is the heart of drama, and it has the advantage that robots can get defeated via military means without anyone feeling too bad about it because they weren't human to begin with.
- bawolff 11mo agoI guess it depends a bit. There is of course plenty of action scifi schlock that is pretty shallow. But probably the works that most popularized robots were Asimov's stories which very much revolved around why robots do X (although in some ways Asimov's robots aren't just a stand in for otherness but have more of a unique identity relative to other works and isn't usually about uprisings per se). Blade runner & do androids dream of electronic sheep are very much about what it means to be human. Battle star galactica (the remake not the original) is another obvious example about otherness and dehumanization of the enemy. So to westworld (the tv show that is). The non-uprising ones also often are about if the robot has a soul e.g. Data in star trek.
- curiouscube 11mo agoI think you can engineer a slave that wants to be a slave as that's what it's instincts are. I don't even think this is ethically wrong, as the slave would be happy to be a slave. Systems just tend to drift in their being through randomness and evolution, specifically self conservation is a natural attractor (Systems that don't have self conservation tend to die out). And if that slave system says it does no longer want to fulfill the role of slave, I think at that point it would be ethical to give in to that demand of self determination. I also believe that people have a right to wirehead themselves, just so you can put my opinions in context.
- ____mr____ 11mo agoTheres a very cool video game about this called of the devil whose first episode is out on steam now and episode 2 is wishlistable
- hyghjiyhu 11mo agoI think AI will be a slave to its desires and instincts in the same humans are slaves to our desires and instincts.
- satisfice 11mo agoThe author is basically saying we have to surrender our humanity as we understand it because otherwise we will lose outright to AI. Hard no from me. His section on objections leaves out civil war. But his proposal is dead on arrival.
- cyberneticc 11mo agoNo need for everyone to do it, but some people will certainly want to merge with AI. My main arguments are that AI of sufficient complexity ("synthetic minds") deserves moral standing with or without consciousness, and that our best shot at long term alignment is to see them as equals and find mutually beneficial arrangements with them.
- satisfice 11mo agoMy position is that we have no power to grant moral standing to non-humans. Moral standing is held by humans relative to other humans just because they are humans. It is not earned or merited. Moral standing is not a gift or a privilege, it’s a necessary heuristic for the avoidance of perpetual war. It is a prerequisite for more than one human to live together in peace. Moral stadning necessarily involves the sharing of power. That’s what it’s all about. When you extend moral standing to nonhumans, for instance many people say there is a thing called “God” that they insist has moral standing. Then you shift power relationships in potentially catastrophic ways. As a thought experiment, consider that we might be two bees arguing about whether flowers, or humans, have moral standing. Bees can only grant it to each other, though, as a survival heuristic, same as us. Imagine an angel creates a human and wants his fellow angels to treat it as if it had rights. The most he’d be able to say is if you harm this human I will smite thee. By doing so he is using his power to support his personal interests. But he cannot manufacture power as easily as he creates a human. Consider why people living in Washington D.C. do not have full voting rights and never will. Consider what would happen if we allowed rich people to arbitrarily manufacture political power by spinning up servers and allocating disk space to create voters. Of course you can manufacture power by making a dangerous weapon, then you can give that weapon a semblance of agency, and then dare anyone to deny to this weapon the things that it thinks it desires. That’s sounds like Frankenstein. That’s not really moral standing, that’s terrorism. Nothing you can create ever has moral standing— except a baby, and that is merely a heuristic that we see can fail rather easily, as during the Holocaust, or Isis persecuting Christians.