5 ms·
Run-and-tumble sounds an awful lot like stochastic gradient descent
by zanecodes 5y ago
Run-and-tumble sounds an awful lot like stochastic gradient descent
- beaconstudios 5y agoyep, the author is describing gradient descent over state-space - which is the driving force behind basically all adaptive systems*, though varying in both state-space complexity and descent algorithm. This should be expected though, as neural net ML practices are a direct descendant of cybernetics, and are essentially an attempt to digitise natural adaptive control systems. * you can reasonably analogise a single sensor/effector control loop as gradient descent against a 1d state.
- lanstin 5y agoDo we assume that searching in real physical space for a literal 3d smooth gradient max corresponds mathematically to maximizing the output of a planned exertion model that is internal to the computing world and probably responding itself to the neurotransmitters?
- beaconstudios 5y agoIf I understand the second half of your comment correctly, then yes - optimising an n-vector fitness value over a state space (for example 2 floats from 0-1 would form a 2d cartesian state space from 0,0 to 1,1 with a continuous output value at all points), is basically the same as finding the tallest sand dune in a desert by walking around, only able to know your current elevation and the angle of the slope underfoot. Of course this is an evaluation where x and z are inputs to our fitness function and y is the output, so this is a 2d space - but this is directly equivalent to say, trying to optimise heat and humidity for maximising yield for a species of plant. You could add further variables like soil acidity or atmospheric CO2 concentration to increase the dimensionality of the state space. I may have some of the exact terminology wrong here (whether 2in-1out is considered 2d or 3d for instance) - my interest is in cybernetics generally not gradient descent specifically - but hopefully you get the gist.
- lanstin 5y agoMy point was that the concentration of the chemicals is not subject to the organisms internal state, and is also basically one dimensional. But my mental model predicting things is subject to the internal state of my dopamine levels, as well as very high dimensional. Pretty sure the math is a lot simpler in the first case. You will get a lot of quasi-periodic and chaotic behavior from the internal model. So bumping up the dopamine a little could easily cause period doubling or quite different, quite non-linear variations. If your gradient descent is walking a high dimensional system with non-linear dynamics, especially in a range where you are getting close to phase changes (e.g. fight or flight), it won’t be a simple little bacterium swimming to the food.
- beaconstudios 5y agoYeah that's where you go from a linear to a nonlinear system, and why the associated discipline is called complexity science. It's still just an extension of the same concept into high dimensional space, but attractors in high dimensional space can be a lot less intuitive than 2d ones.
- lanstin 5y agoSo the mathematics are a lot different and more complicated in the get thru a day as a human case than chemotaxis. Results of stability and convergence of a given control system that work for the linear case can easily fail for the non-linear case. And so while these straightforward maximization systems exist in the human brain, the case for this structure analogized from driving flagella to increase the chemical concentration outside might not be a universal model for human cognition overall.
- cyber_kinetist 5y agoWell, basically all the author did was rediscover reinforcement learning, but with a bit of additional heuristics in how to do epsilon-greedy exploration. I bet 50/50 that the algorithm will work with some tuning and training time, but I doubt what the author said can be extended to anything much more complex than those these simple OpenAI Gym experiments. The problem I have with their argument is a lot more philosophical: it all boils down to the validity of the reward hypothesis (http://incompleteideas.net/rlai.cs.ualberta.ca/RLAI/rewardhypothesis.html http://incompleteideas.net/rlai.cs.ualberta.ca/RLAI/rewardhy...), and how we even define “intelligent” behavior. I’ve written another comment somewhere in this thread for more detail.
- abeppu 5y agoI don't think the author rediscovered reinforcement learning -- they just didn't use that phrase or cite much of the prior work there. But they do mention a Bellman-style decayed sum of rewards, and they use the terms 'state' and 'action' and associated notation which follow the conventions in in RL, so I'm pretty sure they're aware of that field.
- iroh2727 5y agoCorrect. I'm very familiar with RL, but I didn't see how referring to RL literature fit in to my particular argument (and I'm not intimately familiar with that literature). That said, I'd suspect there are fruitful connections to make there, especially considering that this model completely glosses over learning, which is a pretty crucial element... The K&A paper does mention the connection with RL: "Here, we consider the function of the circuit from the point of view of navigation, a connection anticipated by early work on TD learning (Montague et al, 1995)" (specifically, Montague et al found this connection when building a model of bee foraging).