6 ms·
Robust autonomy emerges from self-play
- nine_k 2y agoCan there be "smart toys" for models that help them self-improve in a particularly efficient way?
- grandma_tea 2y agoCan you expand on that? Efficient in what way?
- deleted 2y ago[deleted]
- nine_k 2y agoEfficient in the way of bringing the model to meet the criteria of autonomy faster. On one hand it may be something specifically efficient at reaching some autonomy qualities. OTOH it could be just something that efficiently uses the improvement in the model during training to make the subsequent training faster.
- visarga 2y agoYes, the smart toys are search, code execution, simulations and games.
- jazzyjackson 2y agovideo games are basically like this, progressive level require more skill, learned from the easier levels.
- djmips 2y agoAnd this is a reason we play video games? That they appeal to some ancient instinct to improve?
- krige 2y agoOne of the reasons. I'd wager this is what appeals to people not merely playing but mastering a particular game - playing higher difficulties, 100% completion, and so on. The other reasons would be overcoming other humans (esports/pvp multiplayer), discovery (story driven and exploratory games), and just passing the time (casual games).
- jazzyjackson 2y agoRather, I think it’s borne of necessity to onboard you to the game mechanics. We certainly have a bird brained instinct to catch the worm / win a round, so a good game design cuts you some slack to begin with, so you can have a little dopamine as a treat From there, difficulty should scale up so you don’t always win, giving you that “intermittent reinforcement“ that makes games addictive
- jepj57 2y agoI'd say it's feedback/reward loop plus small, quick to achieve, progressive goal setting.
- Rebuff5007 2y agoIn RL literature this is generally called "curriculum learning". The curriculum is usually modeled as some form of reward function to steer learning, or sometimes by environment configuration (e.g. learn to walk on a normal surface before a slippery surface).
- hirokio123 2y agoI'm creating "smart toys" like that for humans. I recently launched a mobile app. I'd love to see these research breakthroughs feed back into human learning because if humans remain foolish, the world could fall apart. With DeepSeek R1 and these autonomous driving research results, it feels like we've entered an era where human data is no longer necessary. The ability to infinitely expand learning through simulation while maintaining safety in the real world feels like science fiction coming to life—it's truly exciting.
- cainxinth 2y agoA Young Lady's Illustrated Primer
- The28thDuck 2y agoThe concept of being able to simulate 42 years of “experience” in one hour seems so foreign to me. Something about it creeps me out.
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- ThrowawayTestr 2y agoDon't watch the White Christmas Black Mirror episode.
- baq 2y agoMaybe at this point don’t watch any black mirror episodes…?
- ThrowawayTestr 2y agoThis is also an acceptable option. It's good TV but it's also nightmare fuel.
- RGamma 2y agoDon't read Junji Ito's Nagai Yume either.
- geon 2y agoHumanity experiences almost a million years per hour.
- p-a_58213 2y agoIf the gym is sufficiently simple and well-coded, achieving a simulation speed of 367,920x real-time (simulating 42 years in one hour) is plausible. The question is whether these simulated scenarios genuinely reflect 42 years of real-world driving experience and truly represent the information that a single agent has at its disposal when making driving decisions.
- mitthrowaway2 2y agoSomething about dreams that fascinates me is that I usually am genuinely surprised by events that occur in dreams. I interact with other characters whose motivation I cannot understand and whose actions I cannot fully anticipate. It feels like there's a foreign entity acting as DM. This isn't fake surprise. Sometimes I'll wake up and think, "who on earth were those guys and what were they trying to do? And yet their actions make sense..." or, "who came up with that punchline? It's legitimately funny and I never saw it coming, so it can't have been me..." And yet I know it's all being generated by my own brain somehow. Through some kind of privileged access level. And then I think about the bicameral brain structure. Does our brain have two halves so that it can function in a self-play training mode during sleep? Are each halves of my brain experiencing the same dream from opposite points of view? Apologies for the tangent; this is almost totally unrelated to the article and probably something well known to neuroscience for decades. But still, it fascinates me, and the more we learn about the effectiveness of self-play in AI, the more I wonder.
- jes5199 2y agoI think you may have hit upon a novel combination of ideas here. There is something called "social simulation theory" regarding the purpose of dreams, but I don't think it has a neuroanatomical description included.
- HaZeust 2y ago>"I don't think it has a neuroanatomical description included" Genuinely curious, but then why bother?
- jes5199 2y agobecause you can make descriptions of useful human cognition without specifying the implementation. When you’re reading python code, do you ask, “wait, which hard register is this local variable stored in??”
- svnt 2y agoI don’t think this requires two halves, although it certainly seems possible that is what is happening. I believe it only requires that your sensory and post-sensory systems be unpredictably generative when feeding to your subjective sense-making/observer. This could be provided for within a coherent whole brain.
- dang 2y ago[stub for offtopicness]
- awinter-py 2y ago[flagged]
- TZubiri 2y ago[flagged]
- surume 2y ago[flagged]
- dang 2y agoCould you please stop posting unsubstantive comments and flamebait? You've unfortunately been doing it repeatedly. It's not what this site is for, and destroys what it is for. If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
- seaucre 2y agoThis is interesting, and I have always thought this approach worth exploring given the "bitter lesson" in other ML domains, but I think we should be skeptical until we see such models deployed and operating effectively on real-world vehicles.
- dhbradshaw 2y agoInteresting to see this coming out of Apple
- markisus 2y agoSome interesting points from this paper: - All simulated agents use the same neural net with the same weights, albeit with randomized rewards and conditioning vector to allow them to behave as different types of vehicles with different types of aggressiveness. This is like driving in a world where everyone is different copies of you, but some of your copies are in rush while others are patient. This allows backprop to optimize for a sort of global utility across the entire population. - There is no modeling of occlusion effects. Instead, agents are given the state of nearby agents, but corrupted by random noise. In the real world, occluded nearby agents can be extremely close (think about a child running out from behind a parked car). The paper comments on this. > Both Waymax and nuPlan construct observations, maps, and other actors with auto-labeling tools from realworld perception data. This brings occlusion, incorrect or missing traffic-light states, and obstacles revealed at the last moment. Despite the minimalistic noise modeling in GIGAFLOW, the GIGAFLOW policy generalizes zero-shot to these conditions. - The resulting policy simulates agents that are human-like, even though the system has never seen humans drive. This is a great result when one considers other reinforcement learning projects produce extremely high performance agents that humans would consider to be abusive or pathological.
- linux_devil 2y agoMaybe not directly related , I find genertic algorithms and other optimisation algorithms such as Ant Colony Optimisation algorithms intersecting with this approach of self-play and leading to robust autonomy.