4 ms·
It should not take you 5 minutes to make an image classifier in 2022. 30 seconds is a more reasonable amount of time. Dalle is trained using CLIP which you can
by tehsauce 4y ago
It should not take you 5 minutes to make an image classifier in 2022.
30 seconds is a more reasonable amount of time. Dalle is trained using CLIP which you can just use as a zero shot classifier directly, no need to waste time generating images or training a model at all. Just type in the names or descriptions of your classes and your done! Way easier than this :)
- lumost 4y agoI’ve heard such claims for a long time, I can likewise create a classifier out of a simple dice role. It doesn’t say anything about how good it is. Most software applications have low tolerance for error rates, the ones that do have big money being spent on ensuring their accuracy is better than everybody else’s. So while you can make a classifier out of anything in Y time, that doesn’t say anything about whether it’s of any practical use.
- minimaxir 4y agoNo, CLIP is indeed that good. The robustness of its embeddings is the entire reason why VQGAN+CLIP works and can stablely generate images close to the text prompt.
- lumost 4y agoI'm sure it's better than a dice role, but does it beat a modern classifier trained on the domain specific data? EDIT: I raise this issue as over-promises are the death nell for software. Overpromising capability leads to disappointment.
- minimaxir 4y agoSee the zero-shot performance of CLIP: https://openai.com/blog/clip/ https://openai.com/blog/clip/ It's definitely better performance than what you'd get working in 5 minutes from more conventional approaches.
- ShamelessC 4y agoIt tends to, yes. I suggest reading the paper as they discuss this very thing in detail.
- minimaxir 4y agoNormally, this type of comment is Hacker News reductiveness, but yes, image classification via CLIP is that easy, especially with Hugging Face's API for it: https://huggingface.co/docs/transformers/model_doc/clip https://huggingface.co/docs/transformers/model_doc/clip I created a Python package to generate image embeddings from CLIP's vision model without requiring a ML framework (https://github.com/minimaxir/imgbeddings https://github.com/minimaxir/imgbeddings ), and a simple linear classifier on those embeddings does the trick, demo here: https://github.com/minimaxir/imgbeddings/blob/main/examples/cats_dogs.ipynb https://github.com/minimaxir/imgbeddings/blob/main/examples/...
- version_five 4y agoHow big of a model is CLIP? If you're building a phone app that classifies dogs, you may not want to require that it runs some multi-billion parameter monstrosity to perform its comparatively simple task. There is lots of value in building a compact model. "Just" typing in the names ignores the compute you need to have behind the scenes.
- minimaxir 4y agoA compact model is a constraint that changes the problem entirely and doesn't discredit the quick-but-effective approach that works for nearly every other use case.
- tehsauce 4y agoThere are various sizes of CLIP, many are not enormous. For example, one of the base models is just a standard resnet50. So very usable on a mobile device.
- genewitch 4y agoAre you saying that DALL-E is impressive? On fediverse it's used for jokes and memes, because it's really, uh, ugly? simplistic? using obvious components in each image. To my eye, it looks like trickery. Maybe the "full, paid, commercial" model and outputs are better; i'm not sure. I'm actually looking for a decent classifier / object recognition platform to sort on the order of millions of images coarsely - as it stands all of the ones i've tried can't determine if an image is drawn/painted or a photograph, for instance, which reduces my enthusiasm of the whole field. On the other hand, audio AI/ML stuff - such as spleeter - impresses me, as i can't do that stuff by hand.
- ShamelessC 4y agoThe original DALL-E was never released. This is a smaller model made by volunteers. Did you consider using CLIP like parent comment said?
- danlugo92 4y agoDude youve seen dalle 2?