5 ms·
Could you point me toward any more info or keywords on "the cartpoll demo famously tripped up derivative based reinforcement learning for awhile"? It sounds lik
by thatcherc 3y ago
Could you point me toward any more info or keywords on "the cartpoll demo famously tripped up derivative based reinforcement learning for awhile"? It sounds like an interesting bit of history but my google searches are only bringing up python tutorials.
- tangjurine 3y agoChatgpt: What does this refer to: cartpoll demo famously tripped up derivative based reinforcement learning The phrase "cartpole demo famously tripped up derivative-based reinforcement learning" is likely referring to a classic problem in the field of reinforcement learning, which involves balancing a pole on a cart. The pole is attached to the cart via a hinge, and the goal is to keep the pole upright by moving the cart left or right in response to its angle. This problem is often used as a benchmark for testing reinforcement learning algorithms. The phrase suggests that derivative-based reinforcement learning algorithms, which rely on computing gradients of a function with respect to its parameters, were not successful at solving this problem. This could be due to the fact that the problem is highly non-linear and requires precise control, which may be difficult to achieve with gradient-based methods. Edit: bard got it too, with more detail, which is surprising
- littlestymaar 3y agoYou should be really careful with asking this kind of question to ChatGPT, because now you think you've learned the answer, but in fact there are two options with very different outcome: - ChatGPT was trained on a corpus of data containing the answer and is able to give you a decent answer - ChatGPT was never exposed to the answer and will hallucinate a plausible-sounding response, and because it will answer in a really convincing way, you'll get tricked into believing complete bullshit
- Der_Einzige 3y agoSorry, I meant mountain car: https://www.gymlibrary.dev/environments/classic_control/mountain_car/ https://www.gymlibrary.dev/environments/classic_control/moun...
- msackmann 3y agoMy guess: the cart pole is an inverted pendulum, and requires multiple left-right-swing-up movements to bring it from the “hanging” position to the “standing” position. Finding this action sequence using gradients of “where is the tip” vs “where should it be” is very hard, as swinging the pendulum to the left and right goes against the gradient. Instead, using stochastic gradient approximations (policy gradient method such as proximal policy optimization) might be better suited to solving these kinds of problems. Effectively, they do not compute the exact gradient locally, but rather kind of a global approximation by trying out random sequences of actions and determining which of them are closest to the desired outcome. Hence, stochastic gradient approximations might be considered some kind of hybrid between greedy local optimization (such as following the exact gradient) and global optimization.
- Der_Einzige 3y agoSorry, I meant "Mountain Cart" not cartpoll - https://www.gymlibrary.dev/environments/classic_control/mountain_car/ https://www.gymlibrary.dev/environments/classic_control/moun... The reason for this is that the algorithm doesn't like to have to "spend" energy, reducing its score. Without huge amounts of trickery to get the gradient descent algorithm to stop getting stuck in the center, this is never solved - due to using a local optimizer for a global optimization problem (finding good weights in a NN)