3 ms·
>It is not possible to manually label hundreds of millions of images to train a model on them Citation, please? I think you mean "the developers of this techn
by devmor 4y ago
>It is not possible to manually label hundreds of millions of images to train a model on them
Citation, please?
I think you mean "the developers of this technology do not want to pay to have hundreds of millions of images labeled".
- GaggiX 4y agoIt is not believable that someone would pay humans to label 400mln or 5bln images/samples to train a model on them, but I guess if you argument is "everything is possible" then gotcha
- devmor 4y agoYou seem to be confusing "possibility" with your personal opinion on what you think would be done by others.
- GaggiX 4y agoAs a human being I know human limitations, explicitly labeling 400mln/5bln images for a particular task seems absurd to me, but if you think it is realistically possible perhaps you can give an example.
- satvikpendem 4y agoIf it's done in a reCAPTCHA like way, it can be done fairly efficiently and for cheap. In fact Scale AI does just this, they do manual labor operations such as captioning images, as an API. Here's their product for image labeling: https://scale.com/rapid https://scale.com/rapid. Unstable Diffusion is also doing their captioning like how I mentioned, with groups of volunteers as well as hired individuals.
- GaggiX 4y agoScale seems to do, for example, image classification but not captioning as it would be hard to compare the results with others people to verify the quality (when you have a discrete number of classes is really straightforward), also can you report where you read about the Unstable Diffusion plan for manually labeling image datasets? I want to dig deeper
- satvikpendem 4y agoFrom their Reddit post about this: https://old.reddit.com/r/StableDiffusion/comments/zhg18s/unstable_diffusion_here_were_excited_to_announce/ https://old.reddit.com/r/StableDiffusion/comments/zhg18s/uns...
- GaggiX 4y ago> We are releasing Unstable PhotoReal v0.5 trained on thousands of tirelessly hand-captioned images They seem to have created a much smaller dataset than LAION's, it would not work to train a generative model on such a small amount of images (obviously the images here do not have a single domain).
- whiplash451 4y agoThe LAION dataset was designed for the broader community at the first place, so clearly the premise is that they don’t have millions to throw at the problem.