3 ms·
It's pretty clear from the outputs that it's been heavily trained on stock photo libraries. Or maybe that's just the material it extracted the most info from in
by native_samples 4y ago
It's pretty clear from the outputs that it's been heavily trained on stock photo libraries. Or maybe that's just the material it extracted the most info from in the training set. Stock image libraries have enormous numbers of high quality, high res images annotated with simple labels so they're ideal for this type of training task. I'm actually surprised nobody remarked on this yet because the resemblance is so strong.
E.g. look at what it outputs for welcome signs. It can do a great job for that. Many of the images have artistic use of focus, perspective and blur. That's because:
https://www.shutterstock.com/search/photo+of+a+welcome+sign https://www.shutterstock.com/search/photo+of+a+welcome+sign
Likewise the "omg ai bias" examples are easily explainable as being due to training on stock photo libraries:
https://www.shutterstock.com/search/photo+of+a+nurse https://www.shutterstock.com/search/photo+of+a+nurse
Every single output they get for "photo of a <job>" looks exactly like what you'd get from stock photos.
This might also be why it struggles with counting.
https://www.shutterstock.com/search/an+apple https://www.shutterstock.com/search/an+apple
Easy.
https://www.shutterstock.com/search/5+apples https://www.shutterstock.com/search/5+apples
Hmmm ..... not so easy. The search is good: the results that actually have apples in them tend to have five, but, the actual labels don't include the counts.
Likewise with negation. How many images are annotated with "photo of a man NOT running"? My guess is, nearly none.
- cycrutchfield 4y agoHmm my assumption is that they heavily curated the training images and captions with human labelers. I doubt for something like this they would just scrape stock images and use it directly.
- native_samples 4y agoThese models need huge amounts of data. DALL-E was probably trained on the entire web like their other models (I didn't read the paper, maybe they say in there). At any rate the 'stock style' is pretty easy to spot.