4 ms·
The article's use of the word "flaw" is overstated. For background, here are some selected quotes from the article: > "The first part, which you're reading ri
by dj-wonk 8y ago
The article's use of the word "flaw" is overstated.
For background, here are some selected quotes from the article:
> "The first part, which you're reading right now, will set up what RL is and why it is fundamentally flawed."
> "In the typical model of RL, the agent begins only with knowledge of which actions are possible; it knows nothing else about the world, and it's expected to learn the skill solely by interacting with the environment and receiving rewards after every action it takes."
> "how reasonable is it to design AI models based on pure RL if pure RL makes so little intuitive sense?"
To summarize, the article claims that this particular aspect of RL is a "flaw".
I'd suggest it is more useful to call it a design choice. In many cases, this design choice has beneficial properties.
Of course, there are other ways to build learning agents. The field of RL is certainly open to alternatives, including hybrid models and/or relaxing this particular assumption.
I've seen a good number of (popular) articles about RL making rather broad claims, like this article. It appears to me that many of these articles attempt to 'reduce' RL to a smaller/narrower version of itself in order to make their claims. I hope more people start to see that RL is a set of techniques (not a monolith) that can be mixed and matched in many ways for particular applications.
- andreyk 8y agoTo be fair, in the article itself we wind up criticizing "Pure RL" (defined as the basic formulation that is typically followed, in which all learning is done from just the reward signal) and not RL as a whole. We call out a lot of awesome non-pure RL work in the second part and suggest this deserves more attention and excitement over eg AlphaGo.
- dj-wonk 8y agoFair enough. Your article makes a lot of good points, for sure. Here is a quote from the article I want to mention: “Trying to learn the board game 'from scratch' without explanation was absurd, right?” No. It is hardly absurd. Sometimes it works, sometimes not. It is a great starting point, if nothing else. So, I wonder if we have different ideas of what ‘absurd’ means. I agree that we’re in a period of hype. It requires careful work to write clearly without too much zeal or oversimplification. My opinion here is that your attempt to ‘balance’ the debate uses a lot of language that I (and others) perceive as exaggerated.
- jononor 8y agoHas any of the approaches recommended in part 2 been shown to give equivalent or better results than AlphaGo Zero? Eg matching/surpassing human performance with shorter learning time or less compute resources?