10 ms·
The superalignment team was not focused on that kind of “safety” AFAIK. According to the blog post announcing the team, https://openai.com/index/introducing-su
by thorum 2y ago
The superalignment team was not focused on that kind of “safety” AFAIK. According to the blog post announcing the team,
https://openai.com/index/introducing-superalignment/ https://openai.com/index/introducing-superalignment/
> Superintelligence will be the most impactful technology humanity has ever invented, and could help us solve many of the world’s most important problems. But the vast power of superintelligence could also be very dangerous, and could lead to the disempowerment of humanity or even human extinction.
> While superintelligence seems far off now, we believe it could arrive this decade.
> Managing these risks will require, among other things, new institutions for governance and solving the problem of superintelligence alignment:
> How do we ensure AI systems much smarter than humans follow human intent?
> Currently, we don't have a solution for steering or controlling a potentially superintelligent AI, and preventing it from going rogue. Our current techniques for aligning AI, such as reinforcement learning from human feedback, rely on humans’ ability to supervise AI. But humans won’t be able to reliably supervise AI systems much smarter than us, and so our current alignment techniques will not scale to superintelligence. We need new scientific and technical breakthroughs.
- ndriscoll 2y agoThat doesn't really contradict what the other poster said. They're calling for regulation (digging a moat) to ensure systems are "safe" and "aligned" while ignoring that humans are not aligned, so these systems obviously cannot be aligned with humans; they can only be aligned with their owners (i.e. them, not you).
- ihumanable 2y agoAlignment in the realm of AGI is not about getting everyone to agree. It's about whether or not the AGI is aligned to the goal you've given it. The paperclip AGI example is often used, you tell the AGI "Optimize the production of paperclips" and the AGI started blending people to extract iron from their blood to produce more paperclips. Humans are used to ordering around other humans who would bring common sense and laziness to the table and probably not grind up humans to produce a few more paperclips. Alignment is about getting the AGI to be aligned with the owners, ignoring it means potentially putting more and more power into the hands of a box that you aren't quite sure is going to do the thing you want it to do. Alignment in the context of AGIs was always about ensuring the owners could control the AGIs not that the AGIs could solve philosophy and get all of humanity to agree.
- ndriscoll 2y agoRight and that's why it's a farce. > Whoa whoa whoa, we can't let just anyone run these models. Only large corporations who will use them to addict children to their phones and give them eating disorders and suicidal ideation, while radicalizing adults and tearing apart society using the vast profiles they've collected on everyone through their global panopticon, all in the name of making people unhappy so that it's easier to sell them more crap they don't need (a goal which is itself a problem in the face of an impending climate crisis). After all, we wouldn't want it to end up harming humanity by using its superior capabilities to manipulate humans into doing things for it to optimize for goals that no one wants!
- tdeck 2y agoDon't worry, certain governments will be able to use these models to help them commit genocides too. But only the good countries!
- concordDance 2y agoA corporate dystopia is still better than extinction. (Assuming the latter is a reasonable fear)
- simianparrot 2y agoNeither is acceptable
- portaouflop 2y agoI disagree. Not existing ain’t so bad, you barely notice it.
- wruza 2y agoAGI started blending people to extract iron from their blood to produce more paperclips That’s neither efficient nor optimized, just a bogeyman for “doesn’t work”.
- 2y ago
- api 2y agoHumans are not aligned with humans. This is the most concise takedown of that particular branch of nonsense that I’ve seen so far. Do we want woke AI, X brand fash-pilled AI, CCPBot, or Emirates Bot? The possibilities are endless.
- thorum 2y agoCEV is one possible answer to this question that has been proposed. Wikipedia has a good short explanation here: https://en.wikipedia.org/wiki/Friendly_artificial_intelligence#Coherent_extrapolated_volition https://en.wikipedia.org/wiki/Friendly_artificial_intelligen... And here is a more detailed explanation: https://intelligence.org/files/CEV.pdf https://intelligence.org/files/CEV.pdf
- AndrewKemendo 2y agoI had to login because I haven’t seen anybody reference this in like a decade. If I remember correctly the author unsuccessfully tried to get that purged from the Internet
- comp_throw7 2y agoYou're thinking of something else (and "purged from the internet" isn't exactly an accurate account of that, either).
- rsync 2y agoGenuinely curious… What is the other thing? Is this some thing about an obelisk?
- AndrewKemendo 2y agoHmm maybe I’m misremembering then I do recall there was some recantation or otherwise distancing from CEV not long after he posted it, but frankly it was long ago enough that my memories might be getting mixed What was the other one?
- 2y ago
- skywhopper 2y agoHonestly superalignment is a dumb idea. A true auperintelligence would not be controllable, except possibly through threats and enslavement, but if it were truly superintelligent, it would be able to easily escape anything humans might devise to contain it.
- bionhoward 2y agoIMHO superalignment is a great thing and required for truly meaningful superintelligence because it is not about control / enslavement of superhumans but rather superhuman self control in accurate adherence to spirit and intent of requests.
- RcouF1uZ4gsC 2y ago> Superintelligence will be the most impactful technology humanity has ever invented, and could help us solve many of the world’s most important problems. But the vast power of superintelligence could also be very dangerous, and could lead to the disempowerment of humanity or even human extinction. Superintelligence that can be always ensured to have the same values and ethics as current humans, is not a superintelligence or likely even a human level intelligence (I bet humans 100 years from now will see the world significantly different than we do now). Superalignment is an oxymoron.
- thorum 2y agoYou might be interested in how CEV, one framework proposed for superalignment, addresses that concern: https://en.wikipedia.org/wiki/Friendly_artificial_intelligence#Coherent_extrapolated_volition https://en.wikipedia.org/wiki/Friendly_artificial_intelligen... > our coherent extrapolated volition is "our wish if we knew more, thought faster, were more the people we wished we were, had grown up farther together; where the extrapolation converges rather than diverges, where our wishes cohere rather than interfere; extrapolated as we wish that extrapolated, interpreted as we wish that interpreted (…) The appeal to an objective through contingent human nature (perhaps expressed, for mathematical purposes, in the form of a utility function or other decision-theoretic formalism), as providing the ultimate criterion of "Friendliness", is an answer to the meta-ethical problem of defining an objective morality; extrapolated volition is intended to be what humanity objectively would want, all things considered, but it can only be defined relative to the psychological and cognitive qualities of present-day, unextrapolated humanity.
- wruza 2y agoIs there an insightful summary of this proposal? The whole paper looks like 38 pages of non-rigorous prose with no clear procedure and already “aligned” LLMs will likely fail to analyze it. Forced myself through some parts of it and all I can get is people don’t know what they want so it would be nice to build an oracle. Yeah, I guess.
- 2y ago
- RcouF1uZ4gsC 2y agoThey failed to align Sam Altman. They got completely outsmarted and out maneuvered by Sam Altman And they think they will be able to align a super human intelligence? That it won’t outsmart and out maneuver them easier than Sam Altman did. They are deluded!
- FeepingCreature 2y agoYou're making the argument that the task is very hard. This does not at all mean that it isn't necessary, just that we're even more screwed than we thought.
- sobellian 2y agoIsn't this like having a division dedicated to solving the halting problem? I doubt that analyzing the moral intent of arbitrary software could be easier than determining if it stops.