11 ms·
Competitive Self-Play
- Danihan 9y agoThis seems pretty obvious for practicality. The AI can play thousands, or millions of games in different VMs 24/7 and be exposed to a radically higher number of simulated circumstances versus the comparatively plodding rate of genuine interaction.
- aeleos 9y agoThe main issue with self-play is that unless done very methodically it can lead to behavior that does not learn what we want it to learn, but games the simulation and basically cheats. Its not a perfect solution and is just another tool being used to improve models. It can produce really interesting results especially in complex games but its also not perfect.
- vorotato 9y agoOne could argue that "cheating" is a valid solution.
- tlb 9y agoOnly when the learning environment is the same environment it will use when deployed. Agents for robots are often trained in simulation, where it's a common problem for the agent to exploit a physics bug in the simulator.
- recursive 9y agoYou only get what you measure. If you can figure out how to measure the right thing, then that's what you will get.
- gooseus 9y agoI agree... I feel like advanced games will always need guidance to alter the game such that they AI don't reach some local maxima through an exploit. This happens all the time when making games for humans and is evident by how many balancing patches are made to new, highly competitive games (such as any Blizzard title). The next logical step would be for another, impartial AI to observe the games and changing the rules and parameters intelligently as they evolve to guide the player AI toward the actual goal. So, I'll just dive sideways right into a religious/philosophical thought based on the simulation discussion I've been having all over the place: A universe-sized simulation built for a purpose, which requires simulated intelligence to carry out that purpose, would almost certainly include a God intelligence to alter parameters and induce suffering/hardship to direct the simulated intelligence toward that purpose.
- Danihan 9y ago>would almost certainly include a God intelligence to alter parameters and induce suffering/hardship to direct the simulated intelligence toward that purpose. Wow that's pretty deep, actually.
- danohu 9y agoEarlier iterations are buggier and have poorer dev tools. So the God intelligence has more need to smite and command the AIs within the game. After a while the bugs are ironed out, so God can settle back and gently tweak parameters at a distance.
- visarga 9y ago> After a while the bugs are ironed out, so God can settle back and gently tweak parameters at a distance. That explains the hands off approach God has lately with the human society.
- skgoa 9y agoIt is obvious. This is pure PR for OpenAI. They are not doing anything that hasn't been done to death a decade ago.
- dmix 9y agoIt's some excellent PR regardless, no need to be cynical, not to mention highly educational and usually comes with high quality open source code demonstrations with white papers.
- jonny_eh 9y agoI imagine this is the way we get to "true AI", or AI indistinguishable from our own. We train it with a simple virtual environment, that we can gradually increase the complexity of, until it mimics our own. Then we can download the AI into a robot. Boom, it's that easy :P One interesting outcome of this type of AI is that no one knows what the robot's thinking, since no person designed its brain. The brain evolved, just like ours did (but over such a shorter period of real time).
- fizx 9y agoThis is one reason I'm hopeful for a not-killer-robots future. I think there's a chance strong AI will have to evolve over a non-trivial amount of time (as opposed to an afternoon's super-singularity), and a strong evolutionary pressure will be "don't scare the humans."
- deleted 9y ago[deleted]
- mistercow 9y ago"Don't scare the humans" is an insufficient criterion if you're trying to prevent an AI apocalypse. It just means that the AI that destroys us won't be obviously dangerous.
- Twirrim 9y agoGiven sufficient effort, I'm sure an AI out to destroy us could figure out a way to do it without killing us. It very likely even has the advantage of longevity. Harking back to Asimov, and thinking about the way Spacers changed, as they became more and more reliant on robots, it could very well mother us to extinction.
- Filligree 9y agoThat's only a little better. I don't want to go extinct at all.
- eternalcode 9y agoCreate an ICO for it and you'll earn millions /s
- __s 9y agoSo we have a pool of agents, they can send binary blobs, they need to have X many tokens per day to survive (start it low, let it grow over time), they need some way to mine a blockchain, allow some way for them to reproduce via genetic ai & random mutations, add in some misc mechanism of disaster in order to reward saving for a rainy day, maybe add in a mechanism where agents can group up to kill other agents, see if they evolve a means of striking deals & splitting up loot for future attacks Then eventually add a human console where people can invest in agents, see if the agent will give out returns, eventually allowing agents to put smart contracts on ethereum & interact with misc smart contracts..
- gt_ 9y agoI wonder if the AI was as annoyed with the music choice as I was.
- gdb 9y agoAw, that was my one contribution to the video :)! What kind of music would you have preferred instead?
- gt_ 9y agoAs music, I have no problem with it. I just found it distracting for a science video. Turning the sound off helped me to pay attention to the scientific aspects better.
- randyrand 9y agoTo echo the other person, I'd say the music drew too much attention =) Something mellower and more static perhaps? Just tryna give a concrete suggestion.
- gallerdude 9y agoSomething a bit more mellow, it called too much attention to itself.
- PKop 9y agoFirst, turn the volume down. Then music that is quieter, more ambient, and unnoticed, or none at all. Like this[0] or any electronic genre (without lyrics preferable). I also did not like the music. Seems trivial and unnecessary to point out, but that's the point: if people can't help but "notice" the music and view it as a distraction, it was a bad choice. More generally: the cloying, "chipper" stock music of many youtube videos is always irritating to me, and never a good choice for any videos in my opinion. [0] https://youtu.be/85bkCmaOh4o?t=109 https://youtu.be/85bkCmaOh4o?t=109
- tudorw 9y agosome kind of twisted funerary marching tune, something befitting our imminent obsolescence that strikes fear into the heart...
- amelius 9y agoI've seen much better videos of simulated walking structures, e.g. [1]. [1] https://www.youtube.com/watch?v=pgaEE27nsQw https://www.youtube.com/watch?v=pgaEE27nsQw
- ehsankia 9y agoAlso, much better competing AIs (and wittier narration), from back in 1994 [1] [1] https://www.youtube.com/watch?v=JBgG_VSP7f8&t=2m10s https://www.youtube.com/watch?v=JBgG_VSP7f8&t=2m10s
- indescions_2017 9y agoIt seems like fighting games such as Street Fighter or Tekken are a perfect fit for Self-Play. Anyone at OpenAI attempted to build such an agent? Are there any AI research platforms designed specifically for player vs player fighting games? As far as I know, elite human players are still massively dominant. Even though it would make for an exciting matchup. But giving the complexity of actual fighter competition, with combo attacks, power meters, time limits, etc. There is an absurdly high dimension of training variables required. I'd actually like to try and take a step back and apply self-play to something lower dimensional. Perhaps a 2D Tron Light Cycle sim. And see if some truly unexpected strategies arise ;)
- binarymax 9y agoThe research here is more transferable to use in the real world. There is lots of other research in the video game area and it's been going on for awhile [0]. These revelations by OpenAI could potentially be adopted to a physical robot or device. Not sure we can do that with AI Ryu perfecting hadoken :) [0] https://arstechnica.com/gaming/2013/04/this-ai-solves-super-mario-bros-and-other-classic-nes-games/ https://arstechnica.com/gaming/2013/04/this-ai-solves-super-...
- logent 9y agoGiven perfect information, execution, and instant reaction times, fighting games would be trivial for an AI to win at. There are already some bots that will do things like auto-block or parry any incoming attack in Tekken and they're pretty funny to watch (see TOOLASSISTED's videos like this one: https://www.youtube.com/watch?v=nNG5iRMdeg0 https://www.youtube.com/watch?v=nNG5iRMdeg0 ) At high levels of play, where execution largely isn't an issue, it's all about reading your opponent, conditioning them to play how you want to play, and exploiting their tendencies to get in your damage. It would be pretty cool to see a bot learn to play with a human-level reaction time handicap and put them up against a pro in a long set, though!
- bingojess 9y agoFor their dota 2 bot, they included human reaction times if I recall correctly.
- d--b 9y agoMmmh so far, it doesn't look much more compelling than good old genetic algorithms...
- deleted 9y ago[deleted]
- gcb0 9y agoWelcome to the Fad Career! :D Previously relegated to Online Software Engineers, now Data Scientists can feel what it is like to have tons of recent undergrads flooding their fields and coming up with all sorts of "new" ideas that are just the first version of something that is established in the field for decades! Next up, Electrical and Firmware Engineers flooding the IoT Fad Career! Just give it some 5 or 6 years.
- pizza 9y agoMake sure to also look out for epigenetic neuro- noo- quantum cyber- negentropic etc. in the coming years
- noobermin 9y agoThat this isn't higher voted up upsets me. Instead, we see a bunch of circle jerking in the top comments. Somewhat disappointed in HN on this.
- doomjunky 9y agoGenetic algorithms are too slow.
- d--b 9y agoGenetic algorithms didn't have millions of core available for training. It's a different optimization of the same problem.
- cpayne624 9y agoTangentially, I'm so in love with the site design. Shout out to the UI team.
- dag11 9y agoMe too! Mousing over the article links[1] reminds me a lot of the modern Apple TV UI. [1] https://blog.openai.com https://blog.openai.com
- indescions_2017 9y agoLudwig Pettersson https://dribbble.com/luddep https://dribbble.com/luddep
- rtpg 9y agoI'm having a hard time understanding how the body can stay stable at all. For the emergent behavior to appear, you would need the AI to control the body pretty precisely, but if you just had a "random" AI the body would never stay up straight. Seems hard to imagine any amount of generations that get the body up to "stand". I would have maybe expected it to crawl on all fours.
- dweekly 9y ago> Agents initially receive dense rewards for behaviours that aid exploration like standing and moving forward, which are eventually annealed to zero in favor of being rewarded for just winning and losing.
- Nomentatus 9y agoLeading to a thought... are dreams how animals engage in competitive self-play?
- OscarCunningham 9y agoI don't think that can be the only purpose of dreams because I would think that children need to learn a lot more than adults, but they sleep about the same amount.
- koliber 9y agoAt 0:59 in the movie, I noticed another kind of emergent behavior: kicking the goalie where it hurts after it defends the goal. I would call it "retribution".
- kiriakasis 9y agojust a reply to a common reply to this kind of things http://idlewords.com/talks/superintelligence.htm http://idlewords.com/talks/superintelligence.htm
- fil_a_del_fee_a 9y agoImagine a robot police force, based on this technology. If a suspect were to attempt to tackle or evade the police, it would react accordingly. This robot police force would be networked, so all police nationwide would learn from every suspect encounter. You may have seen movies where the human "outsmarts" the robot, but when robots are hundreds of steps ahead of humans, how do we defend ourselves?
- chamoda 9y agoEvolution created consciousness as we experience after billon years of brutal trial and errors. Creation of consciousness will endanger mankind but from evolution point of view, evolution will jump to a next level of evolving. That something evolution could not created directly by herself but her best child mankind created for her.