4 ms·
This sounds like a big deal, can somebody with some expertise in the field comment on it?
by bytefactory 10y ago
This sounds like a big deal, can somebody with some expertise in the field comment on it?
- sherjilozair 10y agoIt has been popular belief that although deep learning has been successful in a lot of domains, its biggest shortcoming is high sample complexity, i.e. even for simple problems you need tens of thousands of samples. Thus, deep learning was believed to be unsuitable for one-shot (or low-shot) learning, where the model needs to learn from one or a handful of samples per class. Brenden Lake et al. showed that a bayesian approach outperforms deep learning based methods in this Nature article: http://web.mit.edu/cocosci/Papers/Science-2015-Lake-1332-8.pdf http://web.mit.edu/cocosci/Papers/Science-2015-Lake-1332-8.p... DeepMind took the challenge, and delivered a deep learning based method, which uses no feature engineering and outperforms human-level performance in low-shot learning, thus providing strong evidence against the widely-held notion that deep learning cannot work with small amounts of data.
- emcq 10y agoDo you have a reference to the work from DeepMind?
- teraflop 10y agohttp://arxiv.org/abs/1605.06065 http://arxiv.org/abs/1605.06065
- colllectorof 10y ago"No feature engineering" claim would be far more convincing if they demonstrated network's abilities on several datasets of different types. It's not like that would be hard to generate synthetically. An interesting test would be "logic puzzles" that use shapes, object counts and so on. For example, a test where you have to label a picture based on the number of objects, regardless of their shapes. As it is, it could be that this particular setup happens to extract features of this particular dataset, while failing miserably on others. --- Thinking about this at a higher level, I ask myself what constitutes true one-shot learning. The reason we care about it is because in real life most problems don't involve huge datasets of available solutions. On the other hand, this paper does involve training a model based on a large dataset of similarly typed, labeled data. The problem the algorithm solves involves items from the same set. The first obvious question is the one I asked above: does this approach work for other types of data? The second one is whether the mechanism would work for more diverse datasets. Finally, the most important question is: how well will it perform on tasks that fall far outside of the initial training data? Because that's the true challenge behind one-shot learning. It's kind of amazing that the paper doesn't try to answer any of those questions. Isn't that the real purpose of research in AI? (Most likely they tried and the results weren't good, but hey, there is no way to tell without re-implementing the whole thing.)