25 ms·
No "Zero-Shot" Without Exponential Data
- cs702 2y agoThis deserves to be on the front page. The authors ask whether image-to-text and text-to-image models (like CLIP and Stable Diffusion) are truly capable of zero-shot generalization. To answer the question, the authors compile a list of 4000+ concepts (see paper for details on how they compile the list of concepts), and test how well 34 different models classify or generate those concepts at different scales of pretraining, from ~3M to ~400M samples. They find that model performance on each concept scales linearly as the concept's frequency in pretraining data grows exponentially, i.e., the rarer the concept the less likely it is actually/properly learned -- which implies there is no "zero-shot" generalization. The authors also release a long-tail test dataset that they cleverly name the "Let it Wag!" benchmark to allow other researchers to see for themselves how current models perform on the long tail of increasingly rare concepts. Go read the whole thing, or at least the introduction. It's clear, concise, and well-written.
- treyd 2y agoBut is this some abstract truth of statistics or is this just a property of how these types of models work?
- kolinko 2y agoWorth noting that this paper is about CLIP only, which is way simpler than llm architectures. (if I’m not mistaken) Still, interesting approach and kind of confirms the experience of most people where clip models can recognize known concepts but struggle with novel ones.
- bilsbie 2y agoHow do we know humans don’t do the same thing?
- tiborsaas 2y agoIt was also mentioned in a just uploaded Computerphile video: https://www.youtube.com/watch?v=dDUC-LqVrPU https://www.youtube.com/watch?v=dDUC-LqVrPU
- deleted 2y ago[deleted]
- stephc_int13 2y agoQuite a few people saw this coming. It is still early to tell if we reached AI winter again or not, but at least we can see that news are slowing down.
- cs702 2y agoIf there are no breakthroughs and funding dries up in the near future, it will feel to many like going off a precipice at high speed. Only the poor souls who survive the fall, at the very bottom of the precipice, will get to experience the AI winter.
- pas 2y agoRAG seems to be all the rage. Not to mention the quest for the cooking up the correct cocktail of smaller MoE/ensemble models, and ... there's decades' worth of optimization work ahead (a few years of it seems to be already VC and edu grants funded), no?
- cs702 2y agoI think the grandparent comment is about AI research, driven by the quest for AGI. Incremental improvement of proven approaches, driven by profit motive, will surely continue regardless of whether there is an AI winter or not.
- pas 2y agoIndeed, but there's probably a very real funding and attention (heh) issue, if there's ton of both then there will be progress. But. Usually the bigger the hype the more progress is expected so the faster the fund-tention will dissipate.
- nkozyra 2y agoRAG takes the current limit of LLM and focuses it on specific problems using custom data. It's not exactly magic, it just finds a way to produce something tangible and usable from an otherwise broadly focused model. It's the rage because it's a way to practically _do work_ from LLMs that generally provide wow from conversationally accurate, often factually accurate responses.
- hprotagonist 2y agoMy long-standing observation has been that while nature may abhor a vacuum, she also really, really loves sigmoids. That performance vs. training data is not linear, but logarithmic, doesn't exactly come as a surprise.
- bilsbie 2y agoEvery exponential is really an s curve?
- hprotagonist 2y agoin physical systems, it’s very often the case!
- mr_mitm 2y agoPretty much always. One exception might be the expansion of the universe
- jskherman 2y agoSymmetries in Nature strike again! It's just like in Noether's theorem.
- eru 2y agoNoether's theorem has nothing to do with S-curves.
- ixaxaar 2y agoThe physical world has limits, that's why sigmoids everywhere.
- CuriouslyC 2y agoThe question that is unanswered, is the logarithmic performance improvement the result of better sampling of the underlying distribution over time, or related to just doing more training with slight variations to effectively regularize the model so it generalizes better? If it's the former, that indicates that we could achieve small models that are every bit as smart as large ones in limited domains, and if that's the case, it radically changes the landscape of what an optimal model architecture is. I suspect from the success of Phi3 that it is in fact the former.
- bearjaws 2y agoComputerphile just did a video on this and it's a pretty good summary: https://www.youtube.com/watch?v=dDUC-LqVrPU https://www.youtube.com/watch?v=dDUC-LqVrPU
- RhysU 2y agoCan one upweight the known-rare concepts in the training set? But then which are rare and how should we identify the rare ones worth knowing? ML is funny. It's predictably going to run into the Education field.
- six_four_eight 2y agoI wonder, is this exponential relation specific to multi-modal models? From my admittedly naïve view it seems to make sense that "...what is rare is not properly learned" would apply generally?
- crote 2y agoThis feels like the worst possible outcome of the current AI hype. We've essentially been ripping off the entire internet and feeding it to the models already, spending many billions of dollars in the process. It's pretty much the largest possible dataset you can currently get, and due to the ever-increasing and now rapidly accelerated AI poisoning of the internet most likely the largest possible dataset which will ever exist. All that and all we're getting out of it is not-entirely-useless but still quite crappy AI? We would've been better off if we had never done this.
- some_random 2y agoQuite crappy? It seems to me like the current SOTA is working plenty good enough for most use cases and like the highest impact (practical) way to improve right now is going to be advancements in domain knowledge acquisition and retention.
- crote 2y agoI sure haven't seen any of those "plenty good" results yet - current SOTA seems to be about as useful as semi-coherently gluing together random Google results. Good enough to perhaps replace a minimum-wage worker, not good enough to provide actual value when you care about the quality of the result. This paper seems to suggest that significant advancement in domain knowledge acquisition and retention is exactly the problem, as you seem to need exponentially more data due to a lack of generalization. What's the point of a model which can perfectly quote Shakespeare if you're a programmer trying to refactor a proprietary codebase and it fails to make a link to whatever garbage it picked up from StackOverflow?
- Filligree 2y agoIn CLIP. Replacing CLIP with an LLM is the current meta in image generation models, specifically because of that lack of generalisation. This isn’t a surprise to anyone.
- greenavocado 2y agoI'm getting a heck of a lot of useful work out of these so-called useless models
- nkozyra 2y agoMy biggest worry is the idea of generating data as training data. We're obviously already unwittingly doing this, but once someone decides to augment low-volume segments of the dataset with generative input, we're going to start getting some really crappy feedback loops.
- nopeNopeNooope 2y ago[dead]
- bilsbie 2y agoDo you ever have novel ideas while walking or in the shower? Well, you’re learning from data you generate.
- nkozyra 2y ago> Well, you’re learning from data you generate. Sure. I'm producing human output from human input in a generally unconstrained, limitless way. This is producing approximated human output from approximated human input. That second level of abstraction will be constrained, ultimately, by the limits of input.
- deleted 2y ago[deleted]
- sottol 2y agoI don't necessarily think it is the equivalent. Maybe it's more akin to a high-schooler reading a single book and then being asked to write several new books that hold equal weight in teaching the next generation of children. These new text books could be great at simplifying the subject matter and making the material accessible or they may just never have fully understood the materials and are misleading. Now imagine that over and over again, imo it's pretty likely to introduce inaccuracies if just taking a naiive approach.
- breck 2y agoWell put. Also, ever read your old journals? You are training on generated data.
- shenberg 2y agoThe CLIP plot (Fig. 2) is damning, however some of the generative models show flat responses in Fig. 3 (e.g. Adobe GigaGAN, DALL-E-mini). While those are on the one hand technically linear relationships, but are also exactly what we'd want: image generation aesthetic score that doesn't care about concept frequency. Maybe the issue is with the contrastive training target used in CLIP?
- twobitshifter 2y agoThe models tested seem to work as expected as it’s not a retrieval model being used here. The weights are lowest on wormsnake but much higher on worm and snake. The temperature of the model for something like stable diffusion has to be higher than something doing a retrieval, so we would not expect it to reproduce the exact worm snake image from its training data.
- xcodevn 2y agoOf course, it will require exponential data for zero shot. The keyword here is zero shot. If you think about it for a second, this applies to humans too. We also need exponential training data to do things without examples.
- IIAOPSW 2y agoWhen we learn the grammar of our language, the teacher does not stand in front of the class and proceed to say a large corpus of examples of ungrammatical sentences, only the correct ones are in the training set. When we learn to drive, we do not need to crash our car a thousand times in a row before we start to get it. When we play a new board game for the first time, we can do it fairly competently (though not as good as experienced players) just by reading and understanding the rules.
- godelski 2y agoI've always been rather upset that it's fairly common to train on things like LAION or COCO and then "zero shot" test on ImageNet. Zero shot doesn't mean a held out set, it means disjoint classes. You can't train on all the animals in the zoo with sentences and then be surprised your model knows zebras. You need to train on horses and test on zebras.
- loandbehold 2y agoHow would the model know what zebra was if it had never seen it? Same is true for humans.
- xboxnolifes 2y agoWhen I was little, zebras were described to me as black and white stripped horses. Without even seeing one, I'm sure anything who has seen a horse could then merge those two concepts to create a close to accurate picture of what a zebra is. If AI is supposed to resemble a human mind with ability to learn, then it must be able to learn from a blanker slate. You don't teach the human before it is born, and in this comparison an AI is born when you finish it's model and set its weights using the training set. If you test it with the training set, you aren't testing ability to comprehend, just regurgitate what it was born with
- dwallin 2y agoIf you trained an image generator, removing all instances of zebras from the training set, you could ask it to output images of a black and white striped horse and it would likely succeed. Then you could fine tune an image recognition model (also with all zebras removed from the training set) on the generated image set to associate it with the word zebra. If you then showed it a bunch of images of actual zebras there’s a really good chance it would succeed.
- salty_biscuits 2y agoYou can look at a medieval bestiary to see how people thought animals might look based on descriptions alone. Like these lovely elephants https://britishlibrary.typepad.co.uk/digitisedmanuscripts/2012/10/elephants-on-parade.html https://britishlibrary.typepad.co.uk/digitisedmanuscripts/20...
- a_wild_dandan 2y agoBetter title: Image classification models suck at identifying nouns that they've rarely seen. Crucial context: - They're only looking at image models -- not LLMs, etc - Their models are tiny - A "concept" here just means "a noun." The authors index images via these nouns. - They didn't control for difficulty in visual representation/recognition of these exceptional infrequent, long-tail "concepts." If I didn't know an object's label, I too would struggle to identify/draw it...
- deleted 2y ago[deleted]
- z7 2y agoOdd how this went from top voted comment to lowest comment. What's the disagreement?
- a_wild_dandan 2y agoI show +23 points. Are popular comments down ranked to encourage diversity? Oh well, I'm proud of the brief success in challenging misinformation here!
- ItsBob 2y agoNot to derail this conversation but... When I'm explaining AI stuff to family, the example I use is classification and I specifically use cats and dogs. I use the analogy of how you teach a toddler that this is a cat and that is a dog. Essentially repetition. And at first they get them mixed up and the parent will say "no, that's a dog" when they think it's a cat and so on. But essentially, for a child learning the difference between a cat and a dog you only need to show them a handful of each and they'll generally get it from that point on. That being said, why does it take ML millions (or billions) of images to be able to say "that's a cat" when a human does it on a handful (might be up to, say 100 but my point stands). Why can ML not do that yet? I'm a dev for many years but not in AI, hence my ELI5 question :) Edit: If the answer is massively long and complicated, perhaps if you could point me to some text (book, paper etc.) and I can read at my leisure. Edit2: I just thought of something. Is it related to whether the child sees a still image or a live cat? So, for example, a still image is a single example of a cat standing in a particular position etc, whereas a moving, live cat, would be interpreted by the brain as many many still images, all processed individually? The end result being that, in fact the child, when seeing a live cat, actually sees thousands or millions of still images of the cat? It just popped into my head there :D
- rvillanueva 2y agoMillions of years of evolution has trained the biological LLM in our brains to be good at fine tuning those concepts.
- stavros 2y agoSaying "it takes a hundred images for a human to learn" implies that you can take a baby/toddler/whatever, who has been blind all their life, restore their sight, show them 100 photos of dogs and cats, and expect them to know what's what. You're ignoring the trillions of frames a toddler has seen before they get to the part where they can even understand what a photo is.
- anon373839 2y agoNeural networks do not imitate the brain and “training” a network has nothing to do with learning, despite these terms serving double duty. Machine learning models are just math functions fitted to some data. When we get predictions from them, we’re really just using a technique to interpolate between the data points. The denser a particular region has been sampled in the data, the better the predictions will be. (This is why GPT-anything will do a good job writing solutions to common leetcode problems, while struggling with a novel problem.) Humans have a powerful abstraction ability far beyond any algorithm that has been developed. We can take in a few pieces of information describing a really unusual set of circumstances, run imaginary experiments and simulations on them, and make very granular and accurate predictions about their consequences. Nobody actually knows how.