3 ms·
The original work. What is the optimization problem being solved precisely? What exactly is done prior to submission to the OpenAI gym? What data does the syste
by highd 9y ago
The original work. What is the optimization problem being solved precisely? What exactly is done prior to submission to the OpenAI gym? What data does the system have access to prior to submission and during runtime?
- RangerScience 9y agoOkay, so, from my understanding: The system has access to available actuators (AFAIK, the X or X+Y position of the agent), a perfect simulator (given this action, that position is the result), and an equation to measure the energy of the system (in a physics / entropy sense). The first example is the inverted pendulum (segway). The agent can move along X, and it takes more energy to go from the down / fallen position to the upright position, than vice versa. Thus, the upright position has better entropy (I never get the +- right with entropy, so I don't know if that means more or less entropy). Since the system knows the entropy present in all possible future states of the system (via the perfect simulator plus the entropy math), it can make a sort of "map", and plot a path from where it is to the global max. In simpler terms, it's optimizing how much energy it takes to get from the current state to all possible future states of the system: in simpler terms, it's way easier (literally, takes less energy) to let the segway fall down than to stand it up in the first place. Does that help?
- highd 9y agoI understand. It appears that constructing the problem this way is a very unfair way to measure if this idea works compared to other reinforcement learning approaches. If you can simulate the system perfectly you can always just simulate k steps for all possible inputs and pick the one that works best.
- RangerScience 9y agoRight, but - how do you measure "what works best"? CEF is an answer to "what works best" that's (theoretically?) applicable to all systems.