5 ms·
Current ~70B models like LLAMA 2 70B are on par wih ChatGPT 3.5. The best smaller models can appear on par at first glance, but they hallucinate at a much highe
by thorum 3y ago
Current ~70B models like LLAMA 2 70B are on par wih ChatGPT 3.5. The best smaller models can appear on par at first glance, but they hallucinate at a much higher rate and lack knowledge of the world. GPT 4 ‘gets’ things at a deeper level and no open source model is even close.
A year is a good timeframe to evaluate things: the rest of the world seems to lag behind OpenAI by around 12-18 months, at least with LLMs and image generation.
On the other hand open source tech usually has additional features for controlling output that OpenAI never bothers to implement, like llama.cpp’s grammars or ControlNet. So in that sense open source is usually ahead of OpenAI in terms of customizability.
- omeze 3y agoI dont think OpenAI is ever going to ahead in image generation, they were lapped very soon after dall-e and every real workflow Ive seen uses Midjourney or Stable Diffusion. The reverse (GPT 4 vision) is well ahead of open source though
- vunderba 3y agoThe original is leagues behind anything current, but DALL-E version 3 absolutely blows any state of the art generative model out of the water, including mid journey 5.2 and SDXL in terms of pure prompt accuracy and coherence. Midjourney still has the edge in quality, but it's a moot point if it takes you 1000 v-rolls to get to your original vision. If all you're generating is anime waifus then MJ/NovelAI/Niji will suffice, but generating prompts particularly featuring relatively complex scenes or actions are amazing on DALL-E 3. And of course unfortunately, it goes without saying that open AI DALL-E is going to be the most restrictive in terms of censorship. I generated these from DALL-E 3 instantly. Try to generate them in any other commercial offering. Go ahead. I'll wait... https://imgur.com/a/2GTRjfK https://imgur.com/a/2GTRjfK Descriptions: A 80s photograph of the Koolaid Man breaking through the Berlin Wall. Comic illustration set at a festive children's party. The main focus is on the magician who looks uncannily like a well-known fictional wizard. He's trying to say abracadabra but accidentally uses the killing curse.
- mattnewton 3y agoSDXL has controlnet for other kinds of non-text input (like scribbles or just masks). The results are much easier to control in my opinion (a picture is worth thousands of prompt words). For pure prompt coherence though I think ideogram is not far behind dalle 3.
- vunderba 3y agoSDXL and even some SD 1.5 checkpoints are great. My current workflow is: 1. Generate initial draft image in DALL-E 3 (iterate as necessary) It's essentially the ONLY good InstructPix2Pix model. 2. Bring into InvokeAI Inpaint with stuff that might be considered censored in DALL-E 3. I'd like to see some proof of Ideogram - it looks... very mobile/instagrammy from the landing page. If you have an account, try out my prompts I'd like to see what you're able to produce.
- vunderba 3y agoEDIT: okay, I just tried Ideogram. It's not terrible and seems to do an okay job on text generation but I'd still say its a distant second compared to DALL-E 3. However, having the ability to maintain image continuity to make refinements of your initial image based on corrections like: "Make the building larger", or "He should have a more prominent forehead" is a game changer (e.g. InstructPix2Pix) and DALL-E 3's the only one that's got it. Ideogram comparisons at bottom: https://imgur.com/a/2GTRjfK https://imgur.com/a/2GTRjfK
- FooBarWidget 3y ago> Midjourney still has the edge in quality, but it's a moot point if it takes you 1000 v-rolls to get to your original vision. I can corroborate this. I wanted about 6 images for a presentation. I rolled ~300 MidJourney images. Most of them looked great, but none of them did what I wanted. I rolled ~50 DALL-E 3 images. In the end, I only picked DALL-E 3 images. They were qualitatively not as good as MidJourney. For example when you zoom in then you see distortions. Or they're a bad fit for 16:9 format. But only DALL-E 3 was able to draw the things I wanted.
- Terretta 3y ago
- kubrickslair 3y agoI have found OpenAI to be the most superior in complex prompts especially where written messages like “Get better, Mom” are expected in the images. The distant second would be ideogram. I am using these tools to send custom personal messages to close friends and family.
- zamadatix 3y agoStrong disagree on this from me as well. DALL-E 3 is miles ahead of the latest Midjourney/Stable Diffusion in image generation. The only real area it falls short vs the other options right now is in how nannying it can be.
- bugglebeetle 3y agoFunction-calling with a JSON schema is about as reliable as llama.cpp’s grammar stuff. I’ve not had any trouble with it.
- infecto 3y agoThe only thing I would argue is that JSON generation and function calling have noticeable decrease in quality of output in certain uses. I have had a hard time writing tests to measure it but its noticeable for my human eyes when I compare various implementations I have written.
- avereveard 3y agoOn the other hand gpt model are converging down. Gpt4 turbo degraded performance so much that now certain 13b produce more consistent results in reasoning. I've a marathon test here for example https://chat.openai.com/share/dfd9b9ae-7214-4dd7-ad20-7ee07abed87c https://chat.openai.com/share/dfd9b9ae-7214-4dd7-ad20-7ee07a... with purposefully open ended and somewhat ambiguous request to see how models perform and gpt4 turbo chat is just not that good it confuses persons out, didn't pick the right one for abduction, didn't change topic when requested, when recalling persons picked the one from the wrong set, when asked to change language it didn't... It know a lot when asked zero shot questions, but when proving it's self consistency and attention it is nowhere near gpt4.
- BoorishBears 3y ago[flagged]
- avereveard 3y agoThe point is exactly that the model people are experiencing is converging down with every subsequent update, and I even mentioned that it's nowhere near the orig gpt4, idk possibly read it again slower instead of jumping to credibility and whatnot.
- BoorishBears 3y ago[flagged]
- avereveard 3y ago> you can't compare it to the "original" via the web ui Good thing then that I was comparing the chat offering to 7b and 13b open models then, maybe you didn't catch that, try a third reading. Also, this kind of is a chatgpt thread. I know it's not a perfect comparison, but still. It's a discussion point I'm presenting, not a research paper.
- 3y ago
- ben_w 3y agoLLMs perhaps (I'm not sure either way, everything moves too quickly), but SDXL 1.0 (July 26, 2023) was a lot better than DALL•E 2 (6 April, 2022). I think DALL•E 3 (August 10, 2023) is a bit better than SDXL, but other than text generation their quality seems very close to me. (That said, perhaps I'm Clever Hands-ing myself by only using SDXL for what it's good at. It's terrible at dragons every time I've tried that…)