3 ms·
I dont think OpenAI is ever going to ahead in image generation, they were lapped very soon after dall-e and every real workflow Ive seen uses Midjourney or Stab
by omeze 3y ago
I dont think OpenAI is ever going to ahead in image generation, they were lapped very soon after dall-e and every real workflow Ive seen uses Midjourney or Stable Diffusion. The reverse (GPT 4 vision) is well ahead of open source though
- vunderba 3y agoThe original is leagues behind anything current, but DALL-E version 3 absolutely blows any state of the art generative model out of the water, including mid journey 5.2 and SDXL in terms of pure prompt accuracy and coherence. Midjourney still has the edge in quality, but it's a moot point if it takes you 1000 v-rolls to get to your original vision. If all you're generating is anime waifus then MJ/NovelAI/Niji will suffice, but generating prompts particularly featuring relatively complex scenes or actions are amazing on DALL-E 3. And of course unfortunately, it goes without saying that open AI DALL-E is going to be the most restrictive in terms of censorship. I generated these from DALL-E 3 instantly. Try to generate them in any other commercial offering. Go ahead. I'll wait... https://imgur.com/a/2GTRjfK https://imgur.com/a/2GTRjfK Descriptions: A 80s photograph of the Koolaid Man breaking through the Berlin Wall. Comic illustration set at a festive children's party. The main focus is on the magician who looks uncannily like a well-known fictional wizard. He's trying to say abracadabra but accidentally uses the killing curse.
- mattnewton 3y agoSDXL has controlnet for other kinds of non-text input (like scribbles or just masks). The results are much easier to control in my opinion (a picture is worth thousands of prompt words). For pure prompt coherence though I think ideogram is not far behind dalle 3.
- vunderba 3y agoSDXL and even some SD 1.5 checkpoints are great. My current workflow is: 1. Generate initial draft image in DALL-E 3 (iterate as necessary) It's essentially the ONLY good InstructPix2Pix model. 2. Bring into InvokeAI Inpaint with stuff that might be considered censored in DALL-E 3. I'd like to see some proof of Ideogram - it looks... very mobile/instagrammy from the landing page. If you have an account, try out my prompts I'd like to see what you're able to produce.
- vunderba 3y agoEDIT: okay, I just tried Ideogram. It's not terrible and seems to do an okay job on text generation but I'd still say its a distant second compared to DALL-E 3. However, having the ability to maintain image continuity to make refinements of your initial image based on corrections like: "Make the building larger", or "He should have a more prominent forehead" is a game changer (e.g. InstructPix2Pix) and DALL-E 3's the only one that's got it. Ideogram comparisons at bottom: https://imgur.com/a/2GTRjfK https://imgur.com/a/2GTRjfK
- FooBarWidget 3y ago> Midjourney still has the edge in quality, but it's a moot point if it takes you 1000 v-rolls to get to your original vision. I can corroborate this. I wanted about 6 images for a presentation. I rolled ~300 MidJourney images. Most of them looked great, but none of them did what I wanted. I rolled ~50 DALL-E 3 images. In the end, I only picked DALL-E 3 images. They were qualitatively not as good as MidJourney. For example when you zoom in then you see distortions. Or they're a bad fit for 16:9 format. But only DALL-E 3 was able to draw the things I wanted.
- Terretta 3y ago> Try to generate them in any other commercial offering. Go ahead. I'll wait... For interest's sake, this is from the second /imagine on MidJourney (so one of the second set of 4 images): https://imgur.com/a/YyWHppb https://imgur.com/a/YyWHppb While yours is what you'd want, this arguably looks more like the super cheesy children's TV commercials back in the day and beats the ideogram take. The Midjourney generations all appear to be referencing Halloween costumes or terrible cosplays, as if there are no trademarked koolaid men in their training set.
- vunderba 3y agoyeah, I did some rolls of this image for MJ but that was back in v4 and wasn't very impressed - doesn't look like its made much progress. The original commercials while silly looking are very visually identifiable as the Koolaid man. I remember hearing that the first versions of MJ used the LAION image set for training data - I'd be curious to see if it has any training data containing the Koolaid man. I did a search through my MJ history from the past year and added the results to the imgur link to include my attempts at generating the Koolaid man from v3/v4/v5.2. https://imgur.com/a/2GTRjfK https://imgur.com/a/2GTRjfK
- Terretta 3y agoIf you limit to the "6+" aesthetic set, there are zero for koolaid man, two for koolaid: https://laion-aesthetic.datasette.io/laion-aesthetic-6pls/images?_search=koolaid&_sort=similarity https://laion-aesthetic.datasette.io/laion-aesthetic-6pls/im... And only three hundred for berlin wall: https://laion-aesthetic.datasette.io/laion-aesthetic-6pls/images?_search=berlin+wall&_sort=similarity https://laion-aesthetic.datasette.io/laion-aesthetic-6pls/im...
- kubrickslair 3y agoI have found OpenAI to be the most superior in complex prompts especially where written messages like “Get better, Mom” are expected in the images. The distant second would be ideogram. I am using these tools to send custom personal messages to close friends and family.
- zamadatix 3y agoStrong disagree on this from me as well. DALL-E 3 is miles ahead of the latest Midjourney/Stable Diffusion in image generation. The only real area it falls short vs the other options right now is in how nannying it can be.