12 ms·
Comparing Adobe Firefly, Dalle-2, and OpenJourney
- famouswaffles 3y agoShould be compared using Bing Image Creator(better version of dall-e) rather than the Dalle-2 site.
- poniko 3y agoMidjurney is still so far ahead it's no competition. Did a lot of testing today and firefly generated so much errors with fingers and stuff, not seen that since the original stability release. Anyone know if the web firefly and the Photoshop version is the same model?
- mettamage 3y agoNot with typography though, haha. It can't spell. I had to draw the letters myself
- jamilton 3y agoNone of these can do text well. There's a model that does do text and composition well, but the name escapes me. And the general quality is much lower overall, so it's a pretty heavy tradeoff.
- kouteiheika 3y agoDeepFloyd? https://github.com/deep-floyd/IF https://github.com/deep-floyd/IF
- throwaway20222 3y agoI believe this is at least one solution, and one that the folks at stability themselves were pushing hard as a next step forward in the development of LLMs.
- deleted 3y ago[deleted]
- jsheard 3y agoIt's worth noting the difference in how the training material is sourced though, Midjourney is using indiscriminate web scrapes while Firefly is taking the conservative approach of only using images that Adobe holds a license for. Midjourney has the Sword of Damocles hanging over its head that depending on how legal precedent shakes out, its output might end up being too tainted for commercial purposes, and Adobe is betting on being the safe alternative during the period of uncertainly and if the hammer does come down on web-scraping models.
- rafark 3y agoWould mid-journey be liable though? I mean you can create copyrighted material using photoshop too. (Even paint!). If I create a Mickey Mouse using photoshop would adobe be liable for it?
- jsheard 3y agoI don't think it really matters whether or not Midjourney themselves are liable, the output of their model being legally radioactive would break their business model either way. They make money by charging users for commercial-use rights to their generations, but a judgement that the generations are uncopyrightable or outright infringing on others copyright would make it effectively useless for the kinds of users who want commercial-use rights.
- suslik 3y agoI wouldn't loose sleep over this if I was working for Midjourney. Copyright lobby is powerful, and when bad actors like Microsoft, Disney etc. jump onto the AI bandwagon and put their legal weight on their side of the lever, everything will turn out well (for them).
- jrm4 3y agoI'm presuming you're not including Stable Diffusion when you say this; the fact that SD and its variants are defacto extremely "free and open source" presently put it way ahead of anything else, and are likely to do so for some time.
- ramraj07 3y agoAs far as I can tell anyone who’s creating images is using midjourney. This is likely the same “Linux is open so it’s way better” tell that to the trillion dollar companies that bet against that.
- GaggiX 3y agoTo be honest most of the AI generated images I find online are generated by Stable Diffusion, the fact that you can't generate NSFW images with MJ makes also a big difference.
- deleted 3y ago[deleted]
- jrm4 3y agoThis comment is breaking my brain. If you're not trolling, like, you do know what operating system the overwhelming vast majority of the "cloud" runs on, yes?
- ramraj07 3y agoI’m perfectly aware of that. But you know what operating systems the overwhelming vast majority of PEOPLE use, yes? Sure likely more machines run Linux on servers but that’s like saying your body has more bacteria than your own cells. Technically correct but actually bullshit.
- jrm4 3y agoAgain, my brain is broken because you mentioned "Trillion dollar COMPANIES," who certainly know the value of Linux, even if a lot of people don't.
- ignite 3y agoIf midjourney could count fingers, I'd be thrilled!
- quitit 3y agoI share the same opinion, but also dislike these tests because each system benefits from a different approach to prompting. What I use to get a good result in MidJourney won't work in StableDiffusion for example. Instead when making these comparisons one needs to set an objective and have people who are familiar with each system to produce their nicest images - since this is a better reflection of the real world usage. For example, ask each participant to read a chapter/page from a book with a lot of specific imagery and then use AI to create what they think that looks like. Regarding image generation in Photoshop I can confirm two things: - It is excellent for in and out painting with a few exceptions* - It remains poor for generating a brand new image *Photoshop's generative fill is very good at extending landscapes, it will match lighting and according to the release video can be smart enough to observe what a reflection should contain even if that is not specifically included in the image (in their launch demo they showed how a reflection pool captured the underside of a vehicle.) Where generative fill falls apart: Inserting new objects that are not well defined produces problems. Choosing something like a VW Beetle will produce a good result as it is well defined, choosing something like "boat", "dragon", or even "pirate's chest": will produce a range of images that do not necessarily fit the scene - this is likely because source imagery for such objects is likely vague and prone to different representations. 1st note about Firefly: Anything that is likely to produce a spherical looking shape tends to be blocked - likely because it resembles certain human anatomy. This is problematic when doing small touch ups such as fixing fingers. A special note about photoshop versus other systems: Photoshop has the added problem of needing to match the resolution of the source material. Currently it achieves this from combining upscaling with resizing - this means that if one is extending an area with high detail, that detail cannot be maintained and instead is softer/blurrier than the original sections. It also means that if one extends directly from the border of an image, then a feathered edge becomes visible which must be corrected by hand. I currently test the following AI generators, feel free to ask me about any of these: StableDiffusion (Automatic and InvokeAI), OpenAI's Dall-E 2, MidJourney, Stability AI's DreamStudio, and Adobe Firefly.
- snowe2010 3y agonot sure this is a good comparison. midjourney likes much shorter prompts, and honestly they're all absolutely terrible for anything that isn't 'photo' based. E.g. ask it to generate a word bubble of the most common programming languages and it will fail every time, no matter what you try. I love it for photo stuff, but for photoshop you'd expect it to be able to do other things as well.
- capybara_2020 3y agoCurious, midjourney does great art and cartoon/comic styles too. Not just realistic images. Most image AI tools are terrible with words. I am curious, what images did you try generating with midjourney?
- jw1224 3y agoThat’s not a fair comparison, as Midjourney is outstanding at a wide range of styles beyond photography. Generating a “word bubble” is going to look terrible in every major diffusion model. Cohesive words and writing in image models is still highly specialised.
- TheOtherHobbes 3y agoIn my first few hours with DiffusionBee I made a couple of very credible semi-abstract portraits by mashing up the styles of unrelated artists. And some splashy watercolours. And some logo line art. And the inevitable booby cheesy rendered forest fairy. I don't think they're terrible at all. They absolutely can make original art with decent production values. They can't write text yet, but I'm sure that's coming soon.
- cainxinth 3y agoAmazing how quickly Dalle-2 went from among the best image transformers to among the worst.
- hathym 3y agochatgpt next...
- denverllc 3y agoWhy innovate when you can regulate?
- flangola7 3y agohttps://time.com/6288245/openai-eu-lobbying-ai-act/ https://time.com/6288245/openai-eu-lobbying-ai-act/
- ralusek 3y agoDall-E 2 was almost immediately displaced by MidJourney. Nothing comes close to even GPT 3.5 at the moment.
- sebzim4500 3y agoAnthropic's models are better than GPT 3.5 in my opinion.
- gwern 3y agoThe stagnation has been very curious. They are part of a large & generally competent org, which otherwise has remained far ahead of the competition, like GPT-4. Except... for DALL-E 2, where it did not just stagnate for over a year (on top of its bizarre blindspots like garbage anime generation), but actually seemed to get worse. They have an experimental model of some sort that some people have access to, but even there, it's nothing to write home about compared to the best models like Parti or eDiff-I etc.
- dvt 3y agoAdobe Firefly is actually extremely competent, especially since it doesn't use copyrighted images in its training set. Using MidJourney (which is fantastic) commercially will be a quagmire for the unlucky company that draws a lawsuit.
- FanaHOVA 3y agoI had done a similar comparison a couple months back but used Lexica instead of DALL-E. Seems clear to me that Midjourney has by far the best "vibes" understanding. Most models get the items right but not the lighting. Firefly seems focused on realism which makes sense for a photography audience. https://twitter.com/fanahova/status/1639325389955952640?s=46&t=IVF1sX_TGndxvax1l-hJ0Q https://twitter.com/fanahova/status/1639325389955952640?s=46...
- rgbrgb 3y agoFor those curious, I tried the same prompts with Kandinsky 2.1 [0]. In my experience it kind of blends the conceptual understanding of DALL-E with the higher quality image generation of Stable Diffusion. Like Midjourney though it kind of injects it's own style and allows you to get "satisfying" results from short prompts. The flaw with these comparisons is that you really shouldn't use the same prompt with different generators. If you want to get best results you do have to play with the prompts and do a bunch of iteration to kind of explore the latent space and find what you're looking for. The first super long prompt looks like it's tuned for stable diffusion for instance. Different generators also have different syntax (e.g. with stable diffusion you can surround a phrase with parens to give it extra emphasis). [0]: https://iterate.world/s/clj4n19u20000jv08iqygiaqw https://iterate.world/s/clj4n19u20000jv08iqygiaqw
- kouteiheika 3y agoFor reference, here's what you can get with a properly tweaked Stable Diffusion, all running locally on my PC. Can be set up on almost any PC with a mid range GPU in a few minutes if you know what you're doing. I didn't do any cherry picking; this is the first thing it generated. 4 images per prompt. 1st prompt: https://i.postimg.cc/T3nZ9bQy/1st.png https://i.postimg.cc/T3nZ9bQy/1st.png 2nd prompt: https://i.postimg.cc/XNFm3dSs/2nd.png https://i.postimg.cc/XNFm3dSs/2nd.png 3rd prompt: https://i.postimg.cc/c1bCyqWR/3rd.png https://i.postimg.cc/c1bCyqWR/3rd.png
- jfdi 3y agoNice! Would you mind sharing which stable diff you used / where you obtained from?
- kouteiheika 3y agoI'm using my own custom trained model. Here, I've uploaded it to civitai: https://civitai.com/models/94176 https://civitai.com/models/94176 There are plenty of other good models too though.
- bavell 3y agoAny tips or guides you followed on training your custom model? I've done a few LoRAs and TI but haven't gotten to my own models yet. Your results look great and I'd love a little insight into how you arrived there and what methods/tools you used.
- kouteiheika 3y agoI'm not an expert at this and there are probably better ways to do this/might not work for you/your mileage may vary, so please take this with a huge grain of salt, but roughly this worked for me: 1. Start with a good base model(s) from which to train from. 2. Have a lot of diverse images. 3. Ideally train for only one epoch. (Having a lot of images helps here.) 4. If you get bad results lower the learning rate and try again. 5. After training try to mix your finetuned model with the original one, in steps of 10%, generate X/Y plot of it, pick the best result. 6. Repeat this process as long as you're getting an improvement. For training I mostly used scripts from here: https://github.com/bmaltais/kohya_ss https://github.com/bmaltais/kohya_ss The main problem here is that essentially during inference you're using a bag of tricks to make the output better (e.g. good negative embeddings), but when training you don't. (And I'm not entirely sure how you'd actually integrate those into the training process; might be possible, but I didn't want to spend too much time on it.) So your fine tuning as-is might improve the output of the model when no tricks are used, but it can also regress it when the tricks are used. Which I why I did the "mix and pick the best one" step. But, again, I'm not an expert at this and just did this for fun. Ultimately there might be better ways to do it.
- mdorazio 3y agoKind of strange to me that they didn't test any prompts with people in them. In my experience that tends to show the limitations of various models pretty quickly.
- usaar333 3y agoLighting also tends to be pretty bad in complex scenes. I find the unrealistic shadows tends to break the photorealism of few light source scenes.
- deleted 3y ago[deleted]
- theobromananda 3y agoAll three of these are horrible, and running Stable Diffusion locally produces incredibly better results as seen in this comment section.
- fumar 3y agoMidJourney produces more consistent and usable results. I am running SD and also pay for MJ. I've tried several checkpoint and loras, but the output is often disappointing or incorrectly using the prompts.
- pdntspa 3y agoWhy didnt this person include Stable Diffusion?
- qiller 3y agoOpenJourney is fine tuned SD
- whatscooking 3y agoI like how simple Firefly’s images are, like something you’d want to work with in Photoshop. Dalle-2 looks terrible. Midjourney is still my favorite.
- chankstein38 3y agoAs someone who has spent hours playing with it in Photoshop (Beta) Firefly is actually pretty damned cool!
- mdorazio 3y agoSince the author didn't have access to Midjourney, here's the first two prompts in MJ with default settings (not upscaled): https://imgur.com/a/siQG06O https://imgur.com/a/siQG06O https://imgur.com/a/vp2oOHu https://imgur.com/a/vp2oOHu
- muhammadusman 3y agothanks for sharing this, do you mind if I include this in the post. I will credit you of course (let me know what you'd like linked to). update: I've edited the post to include these results as well
- throwaway742 3y agoMy result for prompt 2 using Dreamshaper Stable Diffusion model. https://i.imgur.com/ipnf3f5.png https://i.imgur.com/ipnf3f5.png
- soligern 3y ago[flagged]
- abeppu 3y agoIs it intentional that each of the prompts is given twice in that blockquote? It's done without a space, so e.g. in the 2nd example, the word "centeredvalley" appears because of the way the last/first words of the first/second repetition were mashed together. Does that indicate what was actually given to the engines, or was that a copy-paste issue made only while putting together the article? I could imagine that non-words like "cornera" in the last example could throw things off?
- Skywalker13 3y agoAnd here with BlueWillow https://www.bluewillow.ai/ https://www.bluewillow.ai/ 1: https://media.discordapp.net/attachments/1060989219432054835/1120800535306575902/18e3a6c1-5d5d-4947-b677-b220a7cc856d.jpg?width=683&height=683 https://media.discordapp.net/attachments/1060989219432054835... 2: https://media.discordapp.net/attachments/1060989219432054835/1120800584912605184/29cf5372-1f95-4a9a-9d23-9fdea491bcc0.jpg?width=683&height=683 https://media.discordapp.net/attachments/1060989219432054835... 3: https://media.discordapp.net/attachments/1060989219432054835/1120800645616771222/50318f57-de17-49e1-9bd0-a9a226aa9189.jpg?width=683&height=683 https://media.discordapp.net/attachments/1060989219432054835...
- kj_setup 3y agoSeems a lot better than some of the ones in the post
- senko 3y agoFor comparison, these were generated using Stability.ai API: https://postimg.cc/gallery/MQfkgP7/ce388adf https://postimg.cc/gallery/MQfkgP7/ce388adf I used stable-diffusion-xl-beta-v2-2-2 model, copypasted prompts from the blog post, one-shot for each prompt. I chose style presets that closely matched the prompt (added as suffixes in image filenames).
- muhammadusman 3y agoAuthor here: I updated the post to include the generated results from Stable Diffusion and Midjourney (thanks to kouteiheika and mdorazio).
- personjerry 3y agoThe analysis at the end seems to be lacking. From my perspective, PhotoShop and Midjourney come out on top in terms of aesthetic and accuracy, with kouteiheika's Stable Diffusion results[0] a close second. Dall-E falls far behind, which makes sense considering all the work that's gone in to the other systems to fine-tune and build ecosystems around them. [0]: https://news.ycombinator.com/item?id=36408744 https://news.ycombinator.com/item?id=36408744
- SoKamil 3y agoCan we appreciate how well that lightbox works on this site in a mobile mobile browser, especially Safari? Also the gestures are smooth and do not cause any quirks like unintended refresh gesture
- MediumD 3y ago*Shameless Plug* If you want to play around with OpenJourney (or any other fine-tuned StableDiffusion model). I made my own UI with a free tier at https://happyaccidents.ai/ https://happyaccidents.ai/. It supports all open-sourced fine-tuned models & loras and I recently added ControlNet.
- dahwolf 3y agoI'm glad it's not just me getting unusable garbage out of Dall-E and glorious results from MidJourney.
- cubefox 3y agoHere is what the haunted house looks like with Dall-E ~3 (Bing Image Creator): https://www.bing.com/images/create/a-haunted-house-with-ghostly-apparitions2c-eerie-sh/649242295c2c43659807371ae17d875a https://www.bing.com/images/create/a-haunted-house-with-ghos... Generally, this model is much better than Dall-E 2, and it beats Firefly in some areas (I didn't try Midjourney or Stable Diffusion). Firefly usually produces photos with significantly fewer visual mistakes (like the wrong number of fingers or messed up faces) than the Bing Dall-E. But the latter usually understands prompts much better and more often produces something that matches it well. Firefly also doesn't "know" a lot of pop culture or history things, e.g. Marilyn Monroe, or what Coca-Cola is.
- Aeolun 3y ago> small windows opening onto the garden Literally all of the examples have floor to ceiling windows across the entire length of the wall…