4 ms·
I've been trying to parse this body of work - there doesn't seem to be a writeup on the exact implementation, just that it's using "Causal Entropic Forces". The
by highd 9y ago
I've been trying to parse this body of work - there doesn't seem to be a writeup on the exact implementation, just that it's using "Causal Entropic Forces". They have this writeup on the optimization implementation here: https://arxiv.org/pdf/1705.08691.pdf https://arxiv.org/pdf/1705.08691.pdf
One red flag for me is that they're simultaneously claiming that there's no training during the OpenAI Gym while also claiming that the optimization approach is relevant. In that case, what is being optimized? It seems like they might be optimizing over previous simulations - there's frequent reference to having access to a "simulator". In that case, that should effectively count as training, right? I was under the impression that the OpenAI Gym was supposed to benchmark untrained approaches so they could be compared by learning time. Hence the gradually increasing training curves in the other approaches.
- RangerScience 9y agoSo, I'm a little confused by what you're asking - are you asking more about the original ArXiv paper, or the Paulispace post? I think I can explain what you're confused about, if I can understand what it is you're confused about better :)
- highd 9y agoThe original work. What is the optimization problem being solved precisely? What exactly is done prior to submission to the OpenAI gym? What data does the system have access to prior to submission and during runtime?
- RangerScience 9y agoOkay, so, from my understanding: The system has access to available actuators (AFAIK, the X or X+Y position of the agent), a perfect simulator (given this action, that position is the result), and an equation to measure the energy of the system (in a physics / entropy sense). The first example is the inverted pendulum (segway). The agent can move along X, and it takes more energy to go from the down / fallen position to the upright position, than vice versa. Thus, the upright position has better entropy (I never get the +- right with entropy, so I don't know if that means more or less entropy). Since the system knows the entropy present in all possible future states of the system (via the perfect simulator plus the entropy math), it can make a sort of "map", and plot a path from where it is to the global max. In simpler terms, it's optimizing how much energy it takes to get from the current state to all possible future states of the system: in simpler terms, it's way easier (literally, takes less energy) to let the segway fall down than to stand it up in the first place. Does that help?
- highd 9y agoI understand. It appears that constructing the problem this way is a very unfair way to measure if this idea works compared to other reinforcement learning approaches. If you can simulate the system perfectly you can always just simulate k steps for all possible inputs and pick the one that works best.
- RangerScience 9y agoRight, but - how do you measure "what works best"? CEF is an answer to "what works best" that's (theoretically?) applicable to all systems.
- gabrielgoh 9y agoI think an analogy can be made with Bayesian statistics. In principle, Bayesian statistics requires no training, just a way of sampling from the posterior, usually done with expensive MCMC methods. Here, we do not need training of any kind either, just a monte-carlo simulation of the environment and an approximation of which path has the greatest path entropy. Bsaically given a state, you do - Compute the path entropy for all states you can move to - Move into the state with greatest path entropy The tradeoff here is that all the work occurs in inference - every decision requires a complex simulation. In training based approaches the heavy lifting is done during training, and inference is easy
- highd 9y agoYes - the issue is that the work is currently presented as requiring "no training", but it has simply relocated that problem to constructing a perfect simulation of the environment. It then uses the fact that current benchmarking systems have available simulations to "cheat" rather than learning that function itself. One of the most difficult and interesting parts of reinforcement learning is constructing the function that determines the evolution of the system. If you know the evolution function a priori the problem is mostly trivial - i.e. alpha-beta search, graph searching, etc. It's interesting that this merit function works in the absence of a real reward signal, but there's no fair comparison against systems using a reward signal due to this huge alteration to the problem that is providing a perfect simulation.
- gabrielgoh 9y agoi agree completely, and that what's happening is nothing more than brute force search. Though I do think this is still interesting as the reward here is potentially much more well-conditioned than the rewards in RL. Having said that there are situations where this will fail completely, e.g. in maze solving, where the goal is not to play to keep playing but to play to reach the end.
- highd 9y agoIt seems like a more comparable reinforcement learning thing to do would be to combine the entropy criterion with a known reward when available in some way and then do Q learning on that without the simulation requirement. Then in cases where reward is uncertain or infrequent you fall back to a flexibility heuristic.