3 ms·
This is incredibly interesting. Why wasn't this seen as fruitful over the last decade?
by aliljet 3y ago
This is incredibly interesting. Why wasn't this seen as fruitful over the last decade?
- smuser 3y agoMy understanding (not an expert) is a lot of problem domains have very sparse / infrequent rewards - imagine if the only reward you gave a minecraft agent was when it mined a diamond, it would take a lot of gameplay for it to randomly do that and get a reward. So researchers spend time tuning the reward space (oh you mined some dirt, here's a tiny reward. Oh you mined rock, a greater reward, etc) but it's kind of akin to hand crafted feature detection from the pre-neural network days. The Q* mystery is did OpenAI 'solve' reward modelling the same way neural networks solved feature detection.
- pyinstallwoes 3y agoReward for a successful prediction against a goal, then the nuance is defining a goal?
- throwaway4aday 3y agoSounds like the process of tuning the reward space is a type of labelling and ranking problem. If I'm not mistaken, those are two things that GPT-4 is pretty good at. You wouldn't even necessarily pre-label every possible action since GPT-4 could do it in real time.
- eigenvalue 3y agoIt was. Around the time this came out, something like half of the new ML papers were about reinforcement learning. The problem is that it’s incredibly slow and inefficient compared to any learning where you have access to a gradient and can use that to choose more targeted weight updates. But there are certain applications where it’s the only good way of doing it (for example in games, where you don’t have access to a gradient over the space of how good a certain move is given the current game state, and it’s relatively quick and and efficient to simulate the evolution of the game).
- rnimmer 3y agoneuroevolution strategies are another approach to games, for what it's worth (since you said 'only').
- Jagerbizzle 3y agoFor those of us like me who are unfamiliar, can you recommend any useful reading on the topic?
- rnimmer 3y agoYes, take a look at Ken Stanley's web presence: NEAT: https://www.cs.ucf.edu/~kstanley/neat.html https://www.cs.ucf.edu/~kstanley/neat.html HyperNEAT: http://eplex.cs.ucf.edu/hyperNEATpage/HyperNEAT.html http://eplex.cs.ucf.edu/hyperNEATpage/HyperNEAT.html There are explanations, links to the research papers, and links to implementations. Here is a working example that runs in your browser using Javascript: https://liquidcarrot.io/example.flappy-bird/ https://liquidcarrot.io/example.flappy-bird/
- tnecniv 3y agoWell Q learning has been incredibly fruitful over the last decade. If you want to know why it wasn’t fruitful before the last decade, the answer is that it was. Data driven techniques went in and out of vogue for this kind of stuff. The difference this time is how much data we have due to the internet and smart phones and now we have computers capable of drinking from the data fire hose
- cubefox 3y agoBut the recent breakthrough in AI isn't based on Q learning, but on self-supervised learning. Reinforcement learning is only used for part of the fine-tuning of those SSL models.