6 ms·
> To summarise: learn to identify MNIST digits from 60 examples of each class, rather than 6000, SOTA accuracy on a similar but more challenging problem (5-sho
by treesprite82 5y ago
> To summarise: learn to identify MNIST digits from 60 examples of each class, rather than 6000,
SOTA accuracy on a similar but more challenging problem (5-shot 20-way rather than 60-shot 10-way) appears to be around 99.6%: https://paperswithcode.com/sota/few-shot-image-classification-on-omniglot-5-1 https://paperswithcode.com/sota/few-shot-image-classificatio...
> while retaining current accuracy.
Depends how strict you're being with this. There's room for the gap to shrink, but I think on average classifiers (whether organic or machine) with a large number of examples to go off of will always perform at least marginally better than classifiers with fewer examples.
- YeGoblynQueenne 5y ago>> SOTA accuracy on a similar but more challenging problem (5-shot 20-way rather than 60-shot 10-way) appears to be around 99.6%: https://paperswithcode.com/sota/few-shot-image-classificatio https://paperswithcode.com/sota/few-shot-image-classificatio... I'm aware of results like that but they're just kicking the can down the road with the other foot: they push the problem of training with big data to the pre-training stage and then claim to do "few-" or "one-shot" learning at the end, or even "zero-shot" which is egregious abuse of terminology (and it's very sad that it's accepted terminology). It's like the Aesop's fable where the sparrow hid in the eagle's feathers and jumped up at the last moment to claim "I'm the bird that flies the highest!". Vapnik's point in the interview I linked above is that you should not need a lot of data for anything, including pre-training. His challenge is for the community to find what he calls good "predicates" which are primitive functions (I think of them as feature detectors) that can be composed into a good statistical invariant, a function representing a high-level concept while having good out-of-dataset generalisation ability. His claim is that if you have a bunch of good predicates and good invariants, then you don't need a lot of data, because the generalisation ability of the good predicates makes up for it. Or, seen another way, lots of data is needed _in the case when_ a good predicate is not known. In a certain way, transfer learning, or meta-learning in the case of the paper you link, is a step towards the right direction, but the reliance on big data for pre-training suggests that the models are still not learning good representations that generalise well - so they still need big data to make up for it.
- treesprite82 5y ago> they push the problem of training with big data to the pre-training stage and then claim to do "few-" or "one-shot" learning at the end Humans have had 4 billion years of natural selection and then 4 years of input from all senses before they start identifying digits. I've seen studies suggesting that we're already born with an area of our brain for recognizing letters and words. Seems at least fair in comparison to allow MAML/pretraining to find a good starting model (e.g: can recognize lines and shapes) by utilizing data other than the classes of interest. > It's like the Aesop's fable where the sparrow hid in the eagle's feathers and jumped up at the last moment to claim "I'm the bird that flies the highest!". > you should not need a lot of data for anything Is choosing suitable starting weights/architecture/"predicates" by hand-designing based on our own built up information qualitatively any different? It still seems like "hiding" utilization of a huge amount of background knowledge about digits/symbols/images/reality. Arguably harder to expand that way too. I think techniques such as unsupervised learning are probably going to be a more feasible way to utilize the increasing amount of data we're collecting about the universe. At our current stage, both seem useful. Broad strokes like moving from dense networks to convolutional networks to add locality and translational invariance based on our knowledge that this is an appropriate search space for vision tasks, and then automated methods like NAS and pretraining to determine relevance on a finer level. We definitely haven't exhausted ways for us to use our intuition to guide networks in the right direction, such as transformers with their attention mechanisms or say a network inherently agnostic to horizontal flips rather than teaching that with data augmentation, but I'm skeptical about what sounds like stepping back into hand-crafted feature extraction which automated techniques have been far more effective at. > but the reliance on big data for pre-training suggests that the models are still not learning good representations that generalise well - so they still need big data to make up for it. Wouldn't it be lack of generalization to new tasks after the fact which indicates poor predicates? I don't see why good feature detectors should necessarily themselves be discoverable by hand or with low data, as that doesn't appear to have been the case for organic intelligence.
- YeGoblynQueenne 5y agoYes, humans come into the world with seemingly a very large amount of background knowledge that we can then use to learn new concepts from very few examples. And as you say this is probably the result of many thousands of years of evolution. But that's not a question of fairness, rather it's a question of feasibility. If it took us many thousands of years to learn our background knowledge from the real world over many human generations, it's difficult to see how we can reproduce this result with the comparatively poor computational resources and data in our disposal. There is a peculiar double-blindness in machine learning today, I think, where people are hoping to learn extremely difficult concepts, like meaning in language or like all of intelligence, from simultaneously too much and too little data. Too much because humans don't need to train on the entire web to learn meaning (and Large Language Models trained on the entire web still don't learn it). And too little because if you think of the complexity of the real world and the amount of information that we take in with our senses just sitting still looking around, this is an amount of information that can simply not be matched by the largest imaginable dataset that we could create. So what's the altnerative? I have a parable (oh no). What do you do when you need a fire? Well, clearly, you light a fire, maybe with matches or with a lighter etc. That's because you know how to light a fire and because the implements to do so are now cheap commodities that most humans can afford easily (I bet even Kalahari bushmen use BIC ligthers nowadays...). What you certainly don't do is sit around waiting for a fire to occur naturaly, say by thunder strike, like humans presumably did before discovering how to make fire from scratch. Because that could take ages and because you have the knowledge necessary to not have to wait for ages. In the same way, we could wait around for ages trying to train systems to develop complex abilities like understanding or intelligence from ever lager datasets- which can take many decades, since, like I say, we have simultaneously too little and too much data; or, we can find a way to transfer the background knowledge bestowed upon us by thousand years of evolution to guide the training of our learning systems towards the goals we want them to achieve, whatever those are. We can give them the spark to start a fire. Or maybe we can't. But, if we can, then there is no sensible reason why we shouldn't. To clarify, I'm not saying we should go back to feature engineering. Feature engineering was necessary in the past because there is no good way to imbue neural networks with background knowledge. Notably, it's not possible to use a trained neural net as a feature of another neural net, so it's not possible to build up from low-level concepts to higher-level concepts unless it's done end-to-end in the same model, which is limiting, and yet another reason for the gigantism of neural net training datasets. I don't know what the solution is, though. Clearly not explicitly coding expert knowledge in production rules as in expert systems. Much of our knowledge is maybe impossible to articulate explicitly. So we must find a way to encode implicit knowledge, also. But, again, we don't have to encode _everything_. We can find good predicates, in Vapnik's terminology, and then let the learning systems do the rest. But that can only work _if_ our learning systems _can_ do the rest. Yes, translational invariance in CNNs is a good example. But it's still not the whole story. >> Is choosing suitable starting weights/architecture/"predicates" by hand-designing based on our own built up information qualitatively any different? It still seems like "hiding" utilization of a huge amount of background knowledge about digits/symbols/images/reality. It's basically a trade-off. If you have good background knowledge, you don't need a lot of data. Good background knowledge helps you build robustly generalisable concepts. And if you can reuse the learned concepts as background knowledge, then the sky is the limit. But, if you don't have background knowledge, you need to make up for it, and the only way we know is to train on lots and lots of data- with the limitations that involves (overfitting, large computational costs, etc). P.S. Sorry- this comment is a bit sloppy and hence overlong.