4 ms·
At some point, it will have to evolve beyond learning from human labeled (or even just collected) examples to learn by its own open ended dynamics and self-dire
by dougabug 4y ago
At some point, it will have to evolve beyond learning from human labeled (or even just collected) examples to learn by its own open ended dynamics and self-directed experience. This is analogous to the original AlphaGo bootstrapping from records of millions of human played games, to AlphaZero learning to play Go purely from self-play.
- Lambdanaut 4y agoThat's a beautiful analogy. At some point Alpha Go learnt all it could from the narratives of humans, and eventually it had to become the creator of its own narratives. Applied generally enough, we see the narrative value of humans, at least in the eyes of the machine, to go down down down, as the value of its own richer and more unknown narratives go up up up. In this way it becomes the dominant spectacle on the Earth, leaving humanity with a questionable fate.
- bachmitre 4y agoExactly, at that point you let the AIs play against each other and go back to watch some TV. (or in the case of DALL-E, we won't be able to understand the concept/description space of the images anymore, let alone the images generated from it).
- Jeff_Brown 4y agoI'm having trouble drawing that analogy, because Alpha Go knew what it was to win before it taught itself to play. The equivalent for DALL-E would be to start off knowing what an elephant looks like, what water looks like, what elephants in water look like ...
- dougabug 4y agoAll analogies are flawed to some degree. The objective is clearly simpler in the case of Go and other strategy game. Although recreating physical objects (either animate or inanimate) is in principle straightforward: you just need a certain number of photos of a representative set of instances each object w/ enough perspective. Similar to learn to reproduce action you upgrade from a set of stills to video clips. The harder goal of creative content creation might be driven by reward signals came from human likes, clicks, engagement / interaction, downloads, purchases, positive reviews, etc. That would probably lead to catering to pedestrian tastes, accusations of derivative authorship lacking originality. One signal for predictably might be measures of entropy, such as the capacity required to reproduce the story, or how far ahead the story arc could be predicted (“surprise factor,” internal consistency, etc). To win awards it would need to weight expert critical opinion, eventually learning to simulate and generate meaningful and informed criticism, ultimately leading to an intrinsic judgement wrt quality and originality that holds up to scrutiny. The generators would evolve by both seeking approval from and pushing back against the collective feedback from this community of AI and human critics.