4 ms·
This is from memory. Humans wouldn't make this (serious level) of mistakes if they had an enormous data-set of bicycles to draw from.
by dtech 4y ago
This is from memory. Humans wouldn't make this (serious level) of mistakes if they had an enormous data-set of bicycles to draw from.
- jameshart 4y agoThat’s an odd distinction to make. DALL-E is a trained model, everything it produces is ‘from memory’. It doesn’t go out and google for reference material based on the prompt.
- gwern 4y ago> It doesn’t go out and google for reference material based on the prompt. And note that even if it did (which is definitely architecturally possible given progress in retrieval and demos like WebGPT), it still wouldn't necessarily be 'copying', in the goalpost-moving excuse. Facebook has a neat diffusion model which uses retrieval on a dataset of images for generation: https://arxiv.org/abs/2204.02849#facebook https://arxiv.org/abs/2204.02849#facebook You can see the exemplars it uses: they both greatly improve the quality of generated samples, but also don't look that much like the generated sample.
- dtech 4y agoYes, but it was still was trained on a large dataset for the purpose of drawing. It's not an AGI that just happened to be asked to draw bicycles. Maybe a better comparison would be a trained artist, who likely wouldn't make such egregious mistakes.
- gwern 4y agoHumans do, because they have spent a lifetime looking at and riding bikes. They draw from memory just as much as DALL-E 2 does.