4 ms·
For those new to the field of AGI safety: this is an implementation of Paul Christiano's alignment procedure proposal called Iterated Amplification from 6 years
by upwardbound 2y ago
For those new to the field of AGI safety: this is an implementation of Paul Christiano's alignment procedure proposal called Iterated Amplification from 6 years ago. https://www.alignmentforum.org/s/EmDuGeRw749sD3GKd https://www.alignmentforum.org/s/EmDuGeRw749sD3GKd
According to his website he previously ran the language model alignment team at OpenAI. https://paulfchristiano.com/ https://paulfchristiano.com/
It's wonderful to see his idea coming to fruition! I'm honestly a bit skeptical of the idea myself (it's like proposing to stabilize the stack of "turtles all the way down" by adding more turtles - as is insightfully pointed out in this other comment https://news.ycombinator.com/item?id=40817017 https://news.ycombinator.com/item?id=40817017) but every innovative idea is worth a try, in a field as time-critical and urgent as AGI safety.
For a good summary of technical approaches to AGI safety, start with the Future of Life Institute AI Alignment Podcast, especially these two episodes which serve as an overview of the field:
- https://futureoflife.org/podcast/an-overview-of-technical-ai-alignment-with-rohin-shah-part-1/ https://futureoflife.org/podcast/an-overview-of-technical-ai...
- https://futureoflife.org/podcast/an-overview-of-technical-ai-alignment-with-rohin-shah-part-2/ https://futureoflife.org/podcast/an-overview-of-technical-ai...
In both of those episodes, Cristiano's publication series on Iterated Amplification is link #3 in the list of recommended reading.
- pama 2y agoI am not sure why you didnt see the citations to 3 different papers from Cristiano. A simple search in the linked PDF suffices: citations 12, 19, and 31.
- upwardbound 2y agoThank you for providing the reference numbers, this is helpful! I'll update my GP comment.
- TeMPOraL 2y ago> it's like proposing to stabilize the stack of "turtles all the way down" by adding more turtles That's completely fine. Say each layer uses the same amount of turtles, but half the size of the layer above. Even allowing for arbitrarily small turtles, the total height of the infinite stack will be just 2x the height of the first layer. Point being, some series converge to a finite result, including some defined recursively. And in practice, we can usually cut the recursion after first couple steps, as the infinitely long remainder has negligible impact on the final result.
- topherclay 2y agoYou seem to have interpreted the analogy as meaning "you might run out of turtles" instead of something like: "stacks of turtles aren't stable without stable beneath them, no matter how many turtles you use."
- TeMPOraL 2y agoSpace them out. I meant that infinite stack of turtles can be stable, of finite height, and for practical purposes, cut off after few layers without noticeable impact on stability.
- curiousgal 2y ago> AGI safety I genuinely laughed. Oh no somebody please save me from a chatbot that's hallucinating half the time! Joke aside, of course OpenAI is gonna play up how "intelligent" its models are. But it's evident that there's only so much data and compute that you can throw at a machine to make it smart.
- stareatgoats 2y agoThanks, added to my collection of AGI-pessimistic comments that I encounter here, and that I aim to revisit in, say, 20 years. I'm not sure I will be able to say: "you were wrong!". But I do expect so.
- upwardbound 2y agoYou make a good point so I should clarify that Iterated Amplification is an idea that was first proposed as a technique for AGI safety, but happens to also be applicable to LLM safety. I study AGI safety which is why I recognized the technique.
- ben_w 2y agoCovid isn't what most people would call "high intelligence", yet it's a danger because it's heavily optimised for goals that are not our own. Other people using half-baked AI can still kill you, and that doesn't have to be a chatbot as we have current examples from self-driving cars that drive themselves dangerously, and historical examples of the NATO early warning radars giving a false alarm from the moon and the soviet early warning satellites giving false alarms from reflected sunlight, but it can also be a chatbot — there are many ways that this can be deadly if you don't know better: https://news.ycombinator.com/item?id=40724283 https://news.ycombinator.com/item?id=40724283 Every software bug is an example of a computer doing exactly what it was told to do, instead of what we meant. AI safety is about bridging the gap between optimising for what we said vs. what we meant, in a less risky manner than if covid — and while I think it doesn't matter much if covid did or didn't come from a lab leak (the potential that it did means there's an opportunity to improve bio safety there as well as in wet markets), every AI you can use is essentially a continuous supply of the mystery magic box before we know what the word "safe" even means in this context.