3 ms·
Well, not quite. The linked hf space uses the ViT-H-14 OpenCLIP model, which was trained on the Laion-2B dataset[0], which I'd categorize as fitting the reports
by Jackson__ 3y ago
Well, not quite. The linked hf space uses the ViT-H-14 OpenCLIP model, which was trained on the Laion-2B dataset[0], which I'd categorize as fitting the reports description of "noisy and inaccurate image captions" perfectly.
[0] https://laion.ai/blog/large-openclip/ https://laion.ai/blog/large-openclip/
- aargh_aargh 3y agoI see, so they relabelled a smaller (more representative) subset for the captioner manually but more diligently, then used that to relabel their large set (analogous to LAION) with more descriptive captions, then trained DALL-E 3 on that. By the way, in the case of CLIP-Interrogator-2, I wonder how they came up with their wordlists (included in the repo). Are they just all unique terms from OpenCLIP?