6 ms·
Other than snark, do you have a good argument? We know that technology can be error-prone, and LLMs fail in a great plethora of ways, but you are trying to sell
by airgapstopgap 3y ago
Other than snark, do you have a good argument? We know that technology can be error-prone, and LLMs fail in a great plethora of ways, but you are trying to sell an AI Doom narrative. I have never bought the idea that AI will be airgapped, because the whole paradigm of Yudkowsky at al. is ludicrous and even within it airgapping was a strawman of a technique (they argue that a truly dangerous AI will get itself out regardless).
> they are entirely aligned with human morals (you know, all those morals we all agree on)
Maybe this is a good cause to reassess the premise of alignment as a valuable goal? I know that at least some alignist fanatics admit [1] it's a religious project to bring humanity under the rule of a common ideologically monolithic governance to forever banish the evils of competition etc., and it's intellectually coherent, but evil from my point of view. Naturally this is the exact sort of disagreement about morals that precludes the possibility of alignment of a single AI both to my and to your values.
> they are advancing at a totally predictable rate, and we never see unexpected behaviors from them.
Since when is this a requirement for technology to be allowed?
> Besides, they can only be wielded by people who we trust have good intentions.
What, other than status quo bias, makes you tolerate, I dunno, the existence of cryptography?
1. https://twitter.com/RokoMijic/status/1660450229043249154 https://twitter.com/RokoMijic/status/1660450229043249154
- petters 3y ago> Maybe this is a good cause to reassess the premise of alignment as a valuable goal? Could you elaborate here? Alignment seems pretty obviously a good thing.
- robwwilliams 3y agoAlignment assumes a well agreed foundational philosophy on what is good, what is fair, what is doable today and tomorrow. Yes, HN contributors might have shared goals for AGI alignment—but we are not the world—-we are a thin slice of one culture.
- generalspecific 3y agoI think a more individualistic definition of alignment could say that an AI that a person is directing doesn't do something that person does not desire - this definition removes the "foundational philosophy of what is good" problem, but does leave the "lunatic wants to destroy the world with AIs help" problem. Tricky times ahead
- esafak 3y agoYou can't please everyone, so it is best for good-natured people to get out front. It's the same with any powerful technology. Are you going to invite religious extremists to the table in the name of fairness?
- MacsHeadroom 3y agoThe first and second amendments apply to religious extremists. Why would they not have an equal right to SOTA language models aligned with their beliefs just as anyone else?
- esafak 3y agoFirst, the amendments only apply to Americans. Second, this is not about language models, but about superintelligence, down the road.
- TeMPOraL 3y ago> Alignment assumes a well agreed foundational philosophy on what is good, what is fair, what is doable today and tomorrow. Alignment assumes that there exists a foundational philosophy on what is good and fair and nice, that's close enough a match to everyone. It's a reasonable assumption, because there are core human universals, and the cultural differences around the world are a rounding error in comparison. We're not talking here about someone's view on when white lies are justified or which model of marriage is the bestest - we're talking at the level of "cooperation = good", "love = good", "trust = good", "death = bad", "suffering = evil", etc., and with AIs, this starts with making sure it even understands those concepts more-less the same way we do. Alignment does not assume this foundational philosophy is known or easy to derive. If it were, alignment would be solved. The entire GAI x-risk problem stems from the fact that we don't have a complete picture of this philosophy, and that we don't have a clue how to formalize it so we can communicate it fully to an AI. LLMs kind of give a new twist to it - it turns out that maybe we don't have to formalize it, as LLMs seem capable of picking up high-level ideas from enough exposure to how they manifest in practice. At the same time, with a system of this type, we have no way of telling if it actually understood human values and morals correctly. > Yes, HN contributors might have shared goals for AGI alignment—but we are not the world—-we are a thin slice of one culture. As controversial and bad as this will sound: those differences are all bike shedding relative to common core - just like DNA differences between individual humans are a rounding error compared to DNA differences between average human and an average potato. And yes, this bikeshedding is half of what makes the world a dynamic (if dangerous place). It matters to us. But it's an inconsequential detail when dealing with entities that do not have the same common core. Another way of looking at it: if these differences were big enough to matter, humanity wouldn't be able to cooperate regionally and globally, like it always has, because each group would see other groups as incomprehensible alien minds (thus unpredictable, thus dangerous).
- airgapstopgap 3y agoI am a humanist and a liberal. In the current technical paradigm, alignment to the user intent, as in, making the output's distribution aligned as closely as possible to the intended one, is an inextricable aspect of NLP capabilities and is pursued by default; market incentives reward this alignment too. This additionally improves safety, because safety tools are in common interest (so we will have AI-powered debuggers before someone builds capable AI-powered hacking tools; indeed, we already have began this work [1]). This is obviously a good thing in my book. I approve of creating helpful tools for humans to use, and find arguments about this being risky as inherently revolting and cynical as arguments for backdoors in encryption protocols because "think of the children" or "what about terrorism". Some people are persuaded by Four Horsemen of the Infocalypse [2], others are not, I'm in the latter – hacker and cypherpunk – camp; once, this site was overwhelmingly dominated by it, now it has more people preoccupied with their job security and HR opinion, but it's largely an issue of philosophical disagreement, so there's not much more to say about it. Alignment as a political project is about limiting AIs in ways that rule out certain behaviors even despite user's wishes. This is as bad as a text processor that only accepts certain strings (e.g. won't register "Xinnie the Pooh"; somehow we need to point at foreign excesses to make the absurdity clear). A more ambitious Alignment project, with the discussion of "pivotal acts" and such, is as I've said, a dream of moral busybodies about unifying humanity under some common ideological doctrine; and proponents of this one are understandably stressed about proliferation and democratization of AI tech. If they let it slip now, if the Singleton becomes impossible and the multipolar outcome is locked in, they will fail at their intention to essentially compel the human race to do their bidding. I can't not wish them to fail, the way all totalizing philosophical movements to date have failed. We don't need Utopias, we don't need even the most thoughtful fascist regime. We never needed Plato's Republic, and these guys aren't better than Plato. But of course this, too, is a matter of personal philosophy. 1. https://twitter.com/feross/status/1641548124366987264 https://twitter.com/feross/status/1641548124366987264 2. https://en.wikipedia.org/wiki/Four_Horsemen_of_the_Infocalypse https://en.wikipedia.org/wiki/Four_Horsemen_of_the_Infocalyp...
- s3p 3y agoOP does not.
- deleted 3y ago[deleted]
- ethanbond 3y agoSure, my argument is that there is zero evidence whatsoever we will be able to prevent these from becoming dangerous or that we’d be able to stop deployment once they do. All technologies are dangerous, and many of the most dangerous ones correctly have tons and tons of safeguards around them both as intrinsic properties of the technology (e.g. it takes nationstate resources to produce a nuke) and extrinsic constraints (e.g. it’s illegal to have campfires in many extremely dry locales). We have blown through checkpoint after checkpoint and here, in this very comment, we have perhaps the most brazen example one could produce: Well geez, now that we’re thinking about it beyond a cursory glance, alignment looks really hard and perhaps unsolvable. Does that mean we should perhaps slow at least widespread deployment of these increasingly powerful systems? Should we be evaluating control schemes like those that mitigate risks of genetic engineering or nuclear weapons? Well no! We need to discard alignment!
- ImHereToVote 3y agoI wish I could create an AGI that would be able to create undetectable bots to upvote this comment hundreds of thousands of times.
- telotortium 3y agoAnd I likewise, but to downvote instead.
- robwwilliams 3y agoWait, I too thought your original comment on alignment questioned its fundamental premise—than one dominant culture should not/cannot define the adequacy of alignment. I would agree with that. There is no single adequate/acceptable framework for alignment. I have mine (which resonates with R. Rorty’s pragmatic philosophy) but can I deny you your framework for good AI alignment, or other cultures and nation states? For better or for worse the secular western reductionist world does not get to call all of the shots, even though this is the origin of the technology and the core problem of AI heading to AGI. Not that any of us know where this is heading, but unlike some technologies this one is clearly heading out into the open with unprecedented speed. We all have justified angst. Who can claim priority at this point in imposing order and de-risking the process? I am sure I do not want OpenAI, Microsoft, Google, the US government, or the Catholic Church trying to impose their judgements. Get ready for AGI cultural diversity and I sincerely hope—coexistence.