4 ms·
I am new to CNNs/machine learning, but here's my $0.02: Regardless of which technique you use, it seems that the amount of data required to learn is too high. T
by cynicaldevil 10y ago
I am new to CNNs/machine learning, but here's my $0.02:
Regardless of which technique you use, it seems that the amount of data required to learn is too high. This article talks about neural networks accessing billions of photographs, a number which is nowhere near the number of photos/objects/whatever a human sees in a lifetime. Which leads me to the conclusion that we aren't extracting much information from the data. These techniques aren't able to calculate how the same object might look under different lighting conditions, different viewing angles, positions, sizes, and so on. Instead, companies just use millions of images to 'encode' the variations into their networks.
Imo there should be a push towards adapting CNNs to calculate/predict how the object might look under different conditions, which might lead to other improvements. This could also be extended to areas other than image recognition.
- karpathy 10y agoPeople rarely train on billions of images, we're usually around the scale of ~million. This already works quite well in many respects. A back of the envelope calculation assuming about 10fps vision gives ~1B images by age of 5. And humans aren't necessarily starting from scratch as our machine learning systems do. It's not clear if people can calculate what an object might look like in different viewing angles, but even if they could if you would want to in an application, and even if you did there's quite a bit of work on this (e.g. many related papers here http://www.arxiv-sanity.com/1511.06702v1 http://www.arxiv-sanity.com/1511.06702v1). At least so far I'm not aware of convincing results that suggest that doing so improves recognition performance (which in most applications is what people care about).
- cynicaldevil 10y agoEven if we assume that a 5 year old has seen 1000-1500 pictures of say, cats, in his lifespan, it is still far less than the number of images required to train a CNN to label them as accurately as a human can. And of course, I am not talking about just viewing angles. There are several other factors, but I only mentioned the ones which I could think of.
- jm547ster 10y agoEvery second a human opens their eyes, they are seeing a constant stream of changing "pictures" on which to train on.
- kimolas 10y agoThis is the right perspective. It seems the OP believes actual photographs are privileged in some way. In reality, any visual input from our eyes counts as training data, as you said.
- cynicaldevil 10y agoYou seem to forget that the photos are labelled, which counts as supervised learning. What us humans excel at is unsupervised learning, which is difficult for machines. But yes, I agree that humans have the advantage of continuous video access.
- Cybiote 10y agoAnd that posts some advantages we still need to do a lot of work on. Mammals and birds are able to learn online from a few examples per instance, shift to changes in underlying distributions relatively quickly and do so unsupervised.
- karpathy 10y agoA human is very good at one-shot learning but CNNs are actually not too terrible either (and this is also an active area of research, e.g. see http://www.arxiv-sanity.com/1603.05106v2 http://www.arxiv-sanity.com/1603.05106v2). A human might take advantage of good initialization while CNNs start from scratch. Human might have ~1B images by 5 (CNNs get ~1M) of continuous RGBD video and possibly taking advantage of active learning (while CNNs see disconnected samples, which has its pros and cons, mostly cons). i.e. we're disadvantaged in several respects but still doing quite well.
- Cybiote 10y agoIt depends on what we mean by vision. Crows for an example, do the sort of things low level things CNNs are capable of. But for full visual comprehension, they are actively making predictions about physics off a probabilistic world model (learned in part from causal interventions) that feed back into perception. I've not yet looked carefully into it, but I expect that sort of feedback should drastically reduce the amount of required raw data. Machines might not (at first) get to build predictive models from interactions, but even our best approaches to transfer and multi-task learning are very constrained compared to the free form multi modal integrative learning a parrot is capable of. With very little energy spend. This is good, it means there are still a lot of exciting things left to work out.
- Aeolos 10y agoHow many years does it take for a human toddler to be able to form sentences to describe an object he is seeing? We can train a CNN to do that in a few days.
- Cybiote 10y agoThat's not a fair comparison. By that time the toddler can also ask questions, generate new labels using adjectives, label novel instances as compositions of previously acquired knowledge and generate sentences representing complex internal states. They are not limited to observed labels. In fact there is very little supervised learning in the form of [item, label, loss]. Beyond that, with enough stimulation and simply from interacting with each other, children can even spontaneously generate languages with complex grammar; without labeled supervision. They'd also have gained the ability to do very (seriously) difficult things like walking, climbing objects, the rudiments of folk physics, picking things up and throwing things. They'd have some rudimentary ability modeling other agents. It's good to be happy with current progress and I do not suffer from the AI-effect but being too lenient can hamper creativity and impede progress by occluding limitations.
- dbecker 10y agoI've seen many examples where networks are trained on thousands or tens of thousands. The most common example is digit recognition with the MNIST dataset. This is a common problem given to beginners, and even many beginners to CNN's achieve human-level accuracy. That dataset is in the tens of thousands. there should be a push towards adapting CNNs to calculate/predict how the object might look under different conditions Data augmentation like rotations, horizontal flipping and random cropping are a widespread practice.
- cynicaldevil 10y agoI see. Could you link some articles about these techniques?
- dbecker 10y agoI don't have great links for this, but for something less technical, you might look at the blog posts from Kaggle competition winners. Here are a couple examples http://benanne.github.io/2015/03/17/plankton.html http://benanne.github.io/2015/03/17/plankton.html http://blog.kaggle.com/2016/04/13/diagnosing-heart-diseases-with-deep-neural-networks-2nd-place-ira-korshunova/ http://blog.kaggle.com/2016/04/13/diagnosing-heart-diseases-... or check out what's available in a deep learning library like keras http://keras.io/preprocessing/image/ http://keras.io/preprocessing/image/ Sorry I don't have better academic references.