7 ms·
Many concerns in this paper, especially about deep learning being data hungry, limited capacity for transfer, and integrating prior knowledge are addressed by r
by trextrex 8y ago
Many concerns in this paper, especially about deep learning being data hungry, limited capacity for transfer, and integrating prior knowledge are addressed by recent papers in meta-learning and learning-to-learn [1,2,3 and many many others] both in the supervised and reinforcement learning contexts.
In the case of meta-reinforcement learning there has been recent work [5] which seems to indicate that this mechanism is very similar to how learning works in the brain.
In fact, my group recently published a paper [4] about learning-to-learn in spiking networks (making it biologically more realistic) and showing that the network learns priors on families of tasks for supervised learning, and learns useful exploration strategies automatically when doing meta-reinforcement learning.
While I don't claim this is the right path to AGI, it's a very promising and new direction in deep learning research which this paper seems to ignore.
[1] http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.5.323 http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.5.32...
[2] https://arxiv.org/abs/1611.05763 https://arxiv.org/abs/1611.05763
[3] https://arxiv.org/abs/1703.03400 https://arxiv.org/abs/1703.03400
[4] https://arxiv.org/abs/1803.09574 https://arxiv.org/abs/1803.09574
[5] https://www.nature.com/articles/s41593-018-0147-8 https://www.nature.com/articles/s41593-018-0147-8 (Preprint: https://www.biorxiv.org/content/early/2018/04/06/295964 https://www.biorxiv.org/content/early/2018/04/06/295964)
- apsec112 8y agoWhile I haven't read those particular papers yet, one common pattern in the ML literature seems to be: 1) Identify a commonly seen problem with deep learning architectures, whether that's large data volumes, lack of transfer learning, etc. 2) Invent a solution to the problem. 3) Test that the solution works on toy examples, like MNIST, simple block worlds, simulated data, etc. 4) Hint that the technique, now proven to work, will naturally be extended to real data sets very soon, so we should consider the problem basically solved now. Hooray! 5) Return to step #1. If anyone applies the technique to real data sets, they find, of course, that it doesn't generalize well and works only on toy examples. This is simply another form of what happened in the 60s and 70s, when many expected that SHRDLU and ELIZA would rapidly be extended to human-like, general-purpose intelligences, with just a bit of tweaking and a bit more computing power. Of course, that never happened. We still don't have that today, and when we do, I'm sure the architecture will look very very different from 1970s AI (or modern chatbots, for that matter, which are mostly built the same way as ELIZA). I don't mean to be too cynical. Like I said, I haven't read those particular papers yet, so I can't fairly pass judgement on them. I'm just saying that historically, saying problem X "has been addressed" by Y doesn't always mean very much. See also eg. the classic paper "Artificial Intelligence Meets Natural Stupidity": https://dl.acm.org/citation.cfm?id=1045340 https://dl.acm.org/citation.cfm?id=1045340. EDIT: To be clear, I'm not saying that people shouldn't explore new architectures, test new ideas, or write up papers about them, even if they haven't been proven to work yet. That's part of what research is. The problem comes when there's an expectation that an idea about how to solve the problem means that the problem is close to being solved. Most ideas don't work very well and have to be abandoned later. Eg., as one example, this is the Neural Turing Machine paper from a few years back: https://arxiv.org/pdf/1410.5401.pdf https://arxiv.org/pdf/1410.5401.pdf It's a cool idea. I'm glad someone tried it out. But the paper was widely advertised in the mainstream press as being successful, even though it was not tested on "hard" data sets, and (to the best of my knowledge) it still hasn't several years later. That creates unrealistic expectations.
- trextrex 8y agoWhile I don't completely disagree with you, how would you propose researchers go about the problem? If anything, machine learning is applied to real world problems these days more than it ever was. For better or worse, AGI is a hard problem, that's going to take a long time to solve. And we're not going to solve it without exploring what works and what doesn't.
- skywhopper 8y agoI think the mere fact that the OP feels the need to state that (paraphrasing) "possibly additional techniques besides deep learning will be likely necessary to reach AGI" reveals just how deeply the hype has infected the research community. This overblown self-delusion infects reporting on self-driving cars, automatic translation, facial recognition, content generation, and any number of other tasks that have reached the sort-of-works-but-not-really point with deep learning methods. But however rapid recent progress has been, these things won't be "solved" anytime soon, and we keep falling into the trap of believing the hype based on toy results. It'll be better for the researchers, investors, and society to be a little more skeptical of the claim that "computers can solve everything, we're 80% of the way there, just give us more time and money, and don't try to solve the problems any other way while you wait!"
- trextrex 8y agoAgreed. The hype surrounding machine learning is quite disproportionate to what's actually going on. But it's always been that way with machine learning -- maybe because it captures the public's imagination like few other fields do. And there are definitely researchers, top ones no less, who play along with the hype. Very likely to secure more funding, and more attention for themselves and the field. Which has turned out to be quite an effective strategy, if you think about it. The other upside of this hype is that it ends up attracting a lot of really smart people to work on this field, because of the money involved. So each hype cycle leads to greater progress. The crash afterwards might slow things down a bit, particularly in the private sector. But the quantum of government funding available changes much more slowly, and could well last until the next hype cycle starts.