3 ms·
I will contrast with my anecdote.. I build a simulation of 10x10 cityblocks with a population of 30 pedestrians.. I made a reward function based on distance to
by thrax 7y ago
I will contrast with my anecdote.. I build a simulation of 10x10 cityblocks with a population of 30 pedestrians.. I made a reward function based on distance to a random "target point" with penalty for walking in the street vs sidewalk.. and penalty for walking into walls.. the inputs were a 360 degree raycast of 16 samples.. and "distance to target" and the outputs were awsd keyboard inputs. I left it running overnight and by morning, bots were pretty efficiently walking around the city to their random targets via sidewalks. It felt like magic. I could have coded the behavior directly but the learned version seemed somewhat noisier and more organic. This was done a couple years ago using the a JavaScript deepq learning library. It feels like a squishier version of a-star or something.
- basman 7y agoThe raycast probably disambiguated the state pretty well, such that it essentially had to memorize a few hundred actions, so that it did end up basically doing a sort of asynchronous distributed Dijkstra's algorithm.