18 ms·
DeepFloyd IF: open-source text-to-image model
- causality0 3y agoIs this intended to replace Stable Diffusion? Somebody want to give the eli5?
- nickthegreek 3y ago> Stability AI releases DeepFloyd IF, a powerful text-to-image model Hope not. This is a worse license.
- gjsman-1000 3y agoThis is the dumb part about open-source models. Criminals, governments, and propaganda spreaders need not worry about the license; but legitimate users do.
- manv1 3y agoThis is the same problem with laws. The only people that follow them are legitimate users.
- tasty_freeze 3y agoI hear this complaint often, especially in regards to gun control. Yes, there are a subset of people who do what they are going to do irrespective of laws. There is also a middle ground of people where the laws might curtail unwanted behavior. But the main purpose is it provides a basis for punishing unwanted behavior. To make it concrete, one could argue that bank robbers rob banks even though it is illegal, so why have a law against it since law abiding people aren't going to rob banks. Does anyone really think we should remove such laws?
- Jackson__ 3y agoAs far as I can tell from Emad's discord and twitter discussion, the idea appears to be to make this a "research" release, and therefore the worse license. At a later point the model will be renamed "StableIf", and released with a similar license to StableDiffusion.
- mkaic 3y agoYeah Emad was clarifying this on the LAION discord the other day — plan is to have a better-licensed version out Eventually™, guess we'll see how long that takes.
- anticensor 3y agofor certain values of soon™ and eventually™
- dragonwriter 3y agoWell, its a better license than SDXL is available under right now (which is “you can’t have it, but you can use it on StabilityAI’s hosted services”.)
- thejarren 3y agoSecond paragraph in the link: DeepFloyd IF is a state-of-the-art text-to-image model released on a non-commercial, research-permissible license that provides an opportunity for research labs to examine and experiment with advanced text-to-image generation approaches. In line with other Stability AI models, Stability AI intends to release a DeepFloyd IF model fully open source at a future date.
- mkaic 3y agoThis does outperform Stable Diffusion 2.1, but uses a different architecture and requires more memory and compute. Stable Diffusion runs its denoising process in a compressed "latent space" which is how it was able to be so compute-efficient compared to other diffusion models. It also uses the (relatively) small text encoder from OpenAI's CLIP model to encode user prompts. Both of these optimizations meant that it could run much faster compared to say, DALLE or Imagen, but it didn't follow complicated user prompts especially well and had trouble with things like counting and text-rendering. DeepFloyd IF is based on Google's Imagen model, which has two key differences from Stable Diffusion: (1) it denoises in pixel space instead of a compressed latent space, and (2) it uses a 10x larger pretrained text encoder (T5-XXL-1.1) compared to SD's CLIP encoder. (1) allows it to better render high-frequency details and text, and (2) allows it to understand complex prompts much better. These improvements come at the cost of multiple times more memory usage and compute requirements compared to SD, though. In terms of "will it replace SD?"—in the short term I think yes. But I still think latent diffusion models are the future. For example, Stability is gearing up to release Stable Diffusion XL right now, a larger version of the original SD that does higher fidelity and higher resolution generations. I wouldn't be surprised if it takes the crown back from DeepFloyd when it releases, but I guess we'll have to see.
- deleted 3y ago[deleted]
- lucidrains 3y agotldr: bigger text encoder is better. SD will catch up quickly, as conditioning on a new set of precomputed text embeddings is a trivial change
- mkaic 3y agoI didn't think about that but you're totally right, assuming they have those embeddings cached it would be super easy to retrain SD using them. 11B parameter count is rather unfortunate though tbh, I've never been the biggest fan of "scale is all you need" even though it seems to ring irritatingly true most of the time.
- 3y ago
- reckless 3y agoSeems to be entirely a different approach for diffusion. >DeepFloyd IF works in pixel space. The diffusion is implemented on a pixel level, unlike latent diffusion models (like Stable Diffusion), where latent representations are used.
- srajabi 3y agoWow this does so well on text! The original model struggled a lot, it's impressive to see how far they've come.
- mkaic 3y agoI'm quite curious how much of the improvement on text rendering is from the switch to pixel-space diffusion vs. the switch to a much larger pretrained text encoder. I'm leaning towards the latter, which then raises the question of what happens when you try training Stable Diffusion with T5-XXL-1.1 as the text encoder instead of CLIP — does it gain the ability to do text well?
- FL33TW00D 3y agoCheck out figure 4A from the ImageGen paper: https://arxiv.org/pdf/2205.11487.pdf https://arxiv.org/pdf/2205.11487.pdf From that - I would strongly suspect the answer to your question to be yes.
- mkaic 3y agoAh, yes, this seems to be pretty strong evidence. Thanks for pointing that figure out to me!
- minimaxir 3y agoDeepFloyd IF is effectively the same architecture/text encoder as Imagen (https://imagen.research.google/ https://imagen.research.google/), although that paper doesn't hypothesize why text works out a lot better.
- mkaic 3y agoRight, I'm aware of the Imagen architecture, just curious to see further research determining which aspect of it is responsible for the improved text rendering. EDIT: According to the figure in the Imagen paper FL33TW00D's response referred me to, it looks like the text encoder size is the biggest factor in the improved model performance all-around.
- InfosecIcon 3y ago[dead]
- deleted 3y ago[deleted]
- itslennysfault 3y agoThis could be super cool for logos. I've tried using Stable Diffusion to generate logos and it does pretty good at helping brainstorm, but the text is always gibberish so you can use its idea, but you have to add your own text which basically means creating a logo from scratch using its designs as inspiration.
- alex_sf 3y agoYou can't use this to make logos for any commercial product, and it's not safe to use it for hobby projects either, based on the current model license.
- krisoft 3y ago> You can't use this to make logos for any commercial product Yeah good luck figuring out that a particular logo was generated with this particular model. And if someone does good luck doing anything about it. With this amount of fear one wouldn't dare to cross a road without three layers of bubble wrap, plus written authorisation from a lawyer plus a feasibility study from a traffic engineer.
- alex_sf 3y agoYou've never gone through an acquisition or due diligence, have you?
- minimaxir 3y agoGitHub: https://github.com/deep-floyd/IF https://github.com/deep-floyd/IF Colab Notebook for running the model based on the diffusers library: https://colab.research.google.com/github/huggingface/notebooks/blob/main/diffusers/deepfloyd_if_free_tier_google_colab.ipynb https://colab.research.google.com/github/huggingface/noteboo... Hugging Face Space for testing the model: https://huggingface.co/spaces/DeepFloyd/IF https://huggingface.co/spaces/DeepFloyd/IF Note that the model is substantially more compute-intensive than Stable Diffusion, so it may be slower even though that space is running on an A100.
- mkaic 3y agoHF also wrote a blog post on how you can mess around with the model in a python notebook using their excellent Diffusers library: https://huggingface.co/blog/if https://huggingface.co/blog/if
- minimaxir 3y agoI knew the model would have difficulty fitting into a 16GB VRAM GPU, but "you need to load and unload parts of the model pipeline to/from the GPU" is not a workaround I expected. At that point it's probably better to write a guide on how to set up a VM with a A100 easily instead of trying to fit it into a Colab GPU.
- mirekrusin 3y agoWhat about people having RTX 4090 with 24GB or even dual? Does it run on it?
- deleted 3y ago[deleted]
- bwesz 3y agoI tried the HF Space and it generates images of 64x64 resolution, which are basically useless.
- 3y ago
- mkaic 3y agoI think this model will result in a massive new wave of meme culture. AI's already seen success in memes up to this point, but the ability for readable text to be incorporated into images totally changes the game. Going to be an interesting next few months on the interwebz, that's for sure. Exciting times!
- deleted 3y ago[deleted]
- marvinkennis 3y agoSeeing a lot of text-to-image out there recently. Does anyone know what the current state of the art is on image-to-text? Thinking something similar to Midjourney's /describe command that they added in v5
- mkaic 3y agoWhile it's not publicly available yet, I have strong suspicions that multimodal GPT-4 may actually be SOTA in image-to-text. The examples shown in the Sparks of AGI paper were extremely impressive imo, though of course those are cherry-picked so it's unclear how well the model will perform on non-cherry-picked images.
- jah242 3y agoThis is text + image -> text but pretty cool and still might be of interest to you: https://llava-vl.github.io https://llava-vl.github.io
- marvinkennis 3y agoJust entering "Describe this image" in the chat prompt got me exactly what I was looking for. Thanks!
- alex_sf 3y agoThe current license makes this largely unusable for nearly any purpose. Really disappointing release from SAI.
- deleted 3y ago[deleted]
- edkennedy 3y agoI think this is just for pre-release, and they will release fully licensed for Commercial. It doesn't make sense to have a model like this that can do game changing text and logos... but then not license it for commercial. If they don't, that would be ridiculous.
- danwee 3y agoI would be more interested in image-to-text models. Does someone know of any decent model? I saw the GPT4 demo, and they showed that they do image-to-text... but then that was actually a fake (i.e., the model was interpreting the image filename).
- jkea 3y agoMidJourney has the describe function which is kind of like that. Not sure how decent it actually is
- theRealMe 3y agoCan you provide a source on gpt4 image model being fake? I haven’t heard that before, though I have wondered why I haven’t heard anything about the image part and don’t have access to image processing myself.
- runnerup 3y agoAFAIK, converting an image to a text summary isn't really a thing by itself. The related work would be "visual reasoning" which is the ability to ask things about the image in natural language and get responses back also in natural language. I believe the current SOTA test for NLVR is VQAv2[0] or GQA[1]. 0: https://visualqa.org/ https://visualqa.org/ 1: https://arxiv.org/pdf/1902.09506.pdf https://arxiv.org/pdf/1902.09506.pdf
- minimaxir 3y agoFor a fast-but-less-robust model, you can use a ViT encode/GPT-2 decoder model: https://huggingface.co/nlpconnect/vit-gpt2-image-captioning https://huggingface.co/nlpconnect/vit-gpt2-image-captioning For a more-robust-but-hard-to-run model, you can use BLIP2: https://huggingface.co/Salesforce/blip2-opt-2.7b https://huggingface.co/Salesforce/blip2-opt-2.7b
- grumbel 3y agoCLIP Interrogator[1], which is also build into AUTOMATIC1111, gives quite reasonable results, at least if all you need is a prompt, it can't handle complex interactions: Image: https://i.imgur.com/husplYZ.png https://i.imgur.com/husplYZ.png Output: "a white horse with a sign that says rexel's in space, pixelperfect, inspired by Paul Kelpe, official simpsons movie artwork, alternate album cover, in style of nanospace, by Apelles, pickles, pespective, pop surrealism, ingame, in a space cadet outfit, sifi" [1] https://huggingface.co/spaces/pharma/CLIP-Interrogator https://huggingface.co/spaces/pharma/CLIP-Interrogator
- deleted 3y ago[deleted]
- etaioinshrdlu 3y agoDoes paying Hugggingface to run it on the GPU count as commercial use?
- chinaman425 3y ago[dead]
- Thoreandan 3y ago"Hi! I'm B-19-7, but to everyperson I'm called Floyd." -Planetfall (1983) My first thought on seeing "Floyd" and "IF" together. It looks like a Pink Floyd reference from the About page on https://deepfloyd.ai/ https://deepfloyd.ai/ though.
- dr_kiszonka 3y agoLooks like music generation is on their roadmap. Fun! https://stability.ai/careers?gh_jid=4142190101 https://stability.ai/careers?gh_jid=4142190101
- youssefabdelm 3y agoMeh, results feel hodge podge like a bunch of models were stitched together
- vitorgrs 3y agoTried using right now, and it's way better than Stable Diffusion (be it 1.5, 2.1 or SDXL). But is harder to get a good picture. This fine tuned with a good RLHF will be amazing.
- grungegun 3y agoWhat does this mean? Isn't the quality of a model determined by how easy it is to get a good picture?
- vitorgrs 3y agoNot necessarily. IMO a good model needs to follow your prompt well, and that was my problem with Stable Diffusion. I've been trying to get a good portrait picture with "neon lights" on Stable Diffusion and it is almost impossible. Meanwhile with the new Dall-e, that was possible. The picture specially with SDXL is good, but it doesn't really have neon lights... I tried now similar prompt on deepfloyd and managed to get there!
- grungegun 3y agowould be interesting if you could used deepfloyd first for image composition, then apply stable diffusion after for purely stylistic modifications
- vitorgrs 3y agoDefinitely possible :) I've been doing this with new Dall-e + img2img with Stable Diffusion. Explaining: I created a model of me, and wanted to create some good realistic portrait pictures. First I tried to create a model of me using some of the custom models already exist and the result was bad. Then I tried SD 1.5/2.1... It was better, but couldn't really get some of the prompts make real... Then I tried new Dall-e, saved, and inserted my face with img2img on SD and it worked much better!
- marginalia_nu 3y ago> Gorbachev holding meatball pasta in both hands. 1980s synth futuristic max headroom aesthetic. Neon lights. > Aristotle in ancient greek clothes. Toga. New york, rain, film noir, fog, art deco, neon lights, blade runner sci fi Seems to be holding up recently well with the first promt. Second was only OK.
- dang 3y agoRelated: https://stability.ai/blog/deepfloyd-if-text-to-image-model https://stability.ai/blog/deepfloyd-if-text-to-image-model (via https://news.ycombinator.com/item?id=35743727 https://news.ycombinator.com/item?id=35743727, but we've merged that thread into this earlier one)
- jlsreleaf 3y agoWebsite design main page. Bright vibrant neon colors of the rainbow slimes, slime business, kid attention grabbing, splashes of bright neon colors. Professional looking Website page, high quality resolution 8k
- jlsreleaf 3y agoWebsite design for slime. Professional looking, high-quality, 8k, brightest neon colors of the rainbow slimes, splashes of neon colors in background, kid attention grabbing, eye catching
- anirbanc88 3y agohttps://www.kaggle.com/code/anivana/deepfloyd-if-playground/ https://www.kaggle.com/code/anivana/deepfloyd-if-playground/ I played with some ready prompts here
- bulbosaur123 3y agoWhat are the official and unofficial discords? I found only this one on their subreddit: https://discord.gg/GvsvNrVkk5 https://discord.gg/GvsvNrVkk5
- bicepjai 3y agoI understand, I have a decade old 2 nvidia 1080 to card, can we infer and train IF on them ?
- TheBlapse 3y ago"Imagen free"
- TheBlapse 3y agoCurrently down on hugging face
- zimpenfish 3y ago16GB VRAM minimum is a bit steep. Sadly excludes my 3080 which is annoying because I'd like something better than Stable Diffusion locally.
- TaylorAlexander 3y agoIf you don't mind the power consumption I noticed that older nvidia P6000's (24GB) are pretty cheap on ebay! My 16GB P5000 is pretty handy for this stuff.
- coolspot 3y agoLooks like P6000 24Gb goes for $800-$1200 while you can get superior 3090 24Gb for $800-$1000 .
- CamperBob2 3y ago4090s are only $1600 or so now, for that matter.
- TaylorAlexander 3y agooh! My mistake thanks for letting me know.
- NBJack 3y agoAn M40 24GB is less than $200, if you don't mind the trouble to get it's drivers installed, cooled, etc. It's also important to note your motherboard must support larger VRAM addressing; many older chipsets won't be able to boot with it (i.e. some, perhaps almost all, Zen 1 supporters).
- specproc 3y agoThere's a note which suggests you might be able to get by on lower. My 3060 struggles with SD on the defaults, but works fine with float16. There are multiple ways to speed up the inference time and lower the memory consumption even more with diffusers. To do so, please have a look at the Diffusers docs: Optimizing for inference time [1] Optimizing for low memory during inference [2] [1] https://huggingface.co/docs/diffusers/api/pipelines/if#optimizing-for-speed https://huggingface.co/docs/diffusers/api/pipelines/if#optim... [2] https://huggingface.co/docs/diffusers/api/pipelines/if#optimizing-for-memory\ https://huggingface.co/docs/diffusers/api/pipelines/if#optim...
- lalaithion 3y agoHas anyone tried the Scott Alexander AI bet prompts? 1. A stained glass picture of a woman in a library with a raven on her shoulder with a key in its mouth 2. An oil painting of a man in a factory looking at a cat wearing a top hat 3. A digital art picture of a child riding a llama with a bell on its tail through a desert 4. A 3D render of an astronaut in space holding a fox wearing lipstick 5. Pixel art of a farmer in a cathedral holding a red basketball
- swyx 3y agowhere are these prompts from?
- raddles 3y agoScott Alexander made a bet with those prompts here: https://astralcodexten.substack.com/p/a-guide-to-asking-robots-to-design/comment/6945486 https://astralcodexten.substack.com/p/a-guide-to-asking-robo... And followed up with this article when he won the bet: https://astralcodexten.substack.com/p/i-won-my-three-year-ai-progress-bet https://astralcodexten.substack.com/p/i-won-my-three-year-ai...
- epivosism 3y agoYes, I tried them here on an earlier version of IF: https://twitter.com/eb_french/status/1618354180577714176 https://twitter.com/eb_french/status/1618354180577714176
- epivosism 3y agoI thought it was pretty definitive at the time, but when you look really closely (as Scott's opponent is likely to do), it didn't seem like a clear win yet. But that was 3 months ago, and hopefully DF is even better now.
- 55555 3y agoSo this one can create perfect text in images? If true, that’s insane
- GaggiX 3y agoLDM-400M was already able to generate text (predecessor of Stable Diffusion), thanks to the fact that every token in the text encoder (trained from scratch) was available in the attention layer.
- GaggiX 3y agointeresting there are different models: https://github.com/deep-floyd/IF#-model-zoo- https://github.com/deep-floyd/IF#-model-zoo- I'm also very happy for the release of the two upscaler, I can use them to upscale to result of my small 64x64 DDIM models (maybe with some finetuning).
- orra 3y agoNeither the source code nor the weights are open source... This is actually worse than Stability AI's previous offering, in that regard.
- ilaksh 3y agoThey are technically open source. It's just that the model license prohibits commercial use and the code license prohibits bypassing the filters. So it's kind of worse than closed source in a way because it's like a tease. With no API apparently. Theoretically large companies or rich people might be able to make a licensing agreement.
- rgbrgb 3y ago> model license prohibits commercial use I thought that at first, but I think it only prohibits commercial use that breaks regional copyright or privacy laws.
- yellowapple 3y agoThat's already prohibited by, you know, those very same copyright and privacy laws. Adding those same prohibitions to the license not only makes the software nonfree, but pointlessly does so.
- dragonwriter 3y agoIts not pointless, it means the model licensor has a claim against you, as well as whoever would for violating the referenced laws; it also means, and this is probably more important, that in some juridictions, the model licensor has a better defense against liability for contributory infringement if the licensee infringes. EDIT: That said, it’s unambiguously not open source.
- yellowapple 3y ago> it means the model licensor has a claim against you Right, but to what end? The only reason the licensor should care one way or another is the licensor being held liable for what folks do with the software, in which case... > it also means, and this is probably more important, that in some juridictions, the model licensor has a better defense against liability for contributory infringement if the licensee infringes. Do hardware stores need to demand "thou shalt not use this tool to kill people" to their customers to avoid liability for axe murders under such jurisdictions? Or car manufacturers needing to specify "you will not use this product to run over schoolchildren at crosswalks"? Like, I'm sure such jurisdictions exist, but I somehow doubt license terms in an EULA nobody (except for us nerds) will ever read would be sufficient in such a kangaroo court. (EDIT: also, I'm pretty sure the standard warranty disclaimer in your average FOSS license already covers this, without making the software nonfree in the process)
- kingcharles 3y agoThe examples on the README are extremely compelling; the state of the art has been raised yet again.
- simonw 3y agoIt looks like the model on Hugging Face either hasn't been published yet or was withdrawn. I got this error in their Colab notebook: OSError: DeepFloyd/IF-I-IF-v1.0 is not a local folder and is not a valid model identifier listed on 'https://huggingface.co/models https://huggingface.co/models' If this is a private repository, make sure to pass a token having permission to this repo with `use_auth_token` or log in with `huggingface-cli login` and pass `use_auth_token=True`.
- Zetobal 3y agoYou need to accept the license on the HuggingFace model card.
- lerchmo 3y agoit doesn't seem like they have anything published https://huggingface.co/DeepFloyd https://huggingface.co/DeepFloyd
- thewataccount 3y agoI swear I saw it a few minutes ago but I might be crazy.
- Zetobal 3y agoSame got the weights on gdrive.
- famouswaffles 3y agocould you link them ?
- teelelbrit 3y agoDid they just take down the whole model?
- GaggiX 3y ago
- hunkins 3y agoNew restriction in their License suggests the software can't be modified. "2. All persons obtaining a copy or substantial portion of the Software, a modified version of the Software (or substantial portion thereof), or a derivative work based upon this Software (or substantial portion thereof) must not delete, remove, disable, diminish, or circumvent any inference filters or inference filter mechanisms in the Software, or any portion of the Software that implements any such filters or filter mechanisms."
- thewataccount 3y ago> New restriction in their License suggests the software can't be modified. It can be modified. That just says it can't be modified to bypass their filters.
- deleted 3y ago[deleted]
- GaggiX 3y ago>New restriction in their License suggests the software can't be modified. To remove filters. "Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:"
- oh_sigh 3y agoYou can't remove the filters per the license, but the weights will be available soon and so anyone can just reimplement this code using the weights
- Jackson__ 3y agoThere is a similar license clause for the weights[0] as well, so I'm not sure this would apply unless you write the code and train your model from scratch. [0] https://github.com/deep-floyd/IF/blob/main/LICENSE-MODEL#L54 https://github.com/deep-floyd/IF/blob/main/LICENSE-MODEL#L54
- Taek 3y agoFor anyone who doesn't know, DeepFloyd is a StableDiffusion style image model that more or less replaced CLIP with a full LLM (11b params). The result is that it is much better at responding to more complex prompts. In theory, it is also smarter at learning from its training data.
- GaggiX 3y ago>StableDiffusion style Not really, it's a cascaded diffusion model conditioned on the T5 encoder, there is nothing really in common, unless you mean that using a diffusion model is "SD style".
- tmabraham 3y agoIt isn't like Stable Diffusion, it's more like Google's Imagen model.
- dragonwriter 3y ago> It isn’t like Stable Diffusion, it’s more like Google’s Imagen model. Yeah, it looks exactly (architecturally) like Imagen. Google would be running circles around everyone in Generative AI (maybe OpenAI would still have a better core LLM, maybe, but portfolio-wise) if they simply had the ability to cross the gap between building technologies and writing up research papers on them and actually releasing products.
- connerruhl 3y agoThe full release will be soon! https://twitter.com/EMostaque/status/1651328161148174337 https://twitter.com/EMostaque/status/1651328161148174337
- jacob019 3y agoAny web based front ends yet? I put together a system that runs a variety of web based open source AI image generation and editing tools on Vultr GPU instances. It spins up instances on demand, mounts an NFS filesystem with local caching and a COW layer, spawns the services, proxies the requests, and then spins down idle instances when I'm done. Would love to add this, suppose I could whip something up if none exists.
- ronsor 3y agoIt'll probably be in the Auto1111 WebUI within a week.
- jacob019 3y agoYou think? Automatic1111 is still on pytorch 1.7 and SD1.5
- ronsor 3y agoStable Diffusion 2.x has been supported for a while.
- Der_Einzige 3y agoYup, lots of misinformation in this thread from those who are not in-the-know. Automatic1111 is the defacto main UI for these kind of models. It will be supported there, quite quickly.
- nonbirithm 3y agoAutomatic hasn't been updated for several weeks at this point. Several people are trying to fork the repo to make their own continuation.
- teelelbrit 3y agoWhat's your app / service called?!
- atleastoptimal 3y ago> Text > Hands good god it solves the two biggest meme issues with image models in one go. Will this be the new state of the art every other model is compared to?
- Taek 3y agoThere are good reasons to believe that this will be the new state of the art by a comfortable margin. Hard to know until we can actually play with it.
- gwern 3y agoWe already knew those were going to be solved by scale like using T5 instead of the really small bad text encoder SD used, because they were solved by Imagen etc.
- astrange 3y agoThere's fundamental tradeoffs, as there always will be when you're compressing things into an image model. So, this is going to have new different issues. Since it's similar to Imagen, it probably can't handle long complex prompts as well, since they developed Parti afterward. Here's my question: are there any image models where, if you prompt "1+1", you get an image showing "3"?
- dragonwriter 3y ago> So, this is going to have new different issues. Well, yeah, its a bigger set of models (particular the language model) that takes more resources (both to train and for inference.) That’s the tradeoff. > Here’s my question: are there any image models where, if you prompt “1+1”, you get an image showing “3”? You want a t2i model that does arithmetic in the prompt, translates to it to “text displaying the number <result>”, but, also does the arithmetic wrong? Yeah, I don’t think that combination of features is in any existing model or, really, in any of the datasets used for evaluation, or otherwise on anyone’s roadmap.
- astrange 3y ago
- epivosism 3y agoExample of how much better it can do compared to midjourney, on a complex prompt: https://twitter.com/eb_french/status/1623823175170805760 https://twitter.com/eb_french/status/1623823175170805760 It is able to put people on the left/right and put the correct t-shirts and facial expressions on each one. This is compared to mj which just mixes together a soup of every word you use and plops it out into the image. Huge MJ fan of course, it's amazing, but having compositional power is another step up.
- Zetobal 3y agoThat's not how any of this works what do you even mean by compositional power? Every model speaks a different "language" comparing prompts like this has no merit and shows only that the person who makes the claim lacks understanding of the subject matter.
- epivosism 3y agoCompositional power might mean "the image more resembles the composition you want and describe" i.e. if you say "a red cube on a green sphere" in DeepFloyd, you will get it. If you say that in MJ, you won't. That means you have more power to compose the image you want with this tool.
- Zetobal 3y agoNo, it does mean you don't understand how to prompt MJ, you don't understand it's language. You might like french more but it doesn't mean that it's a better language than english. MJ even says that their model doesn't understand language like humans do in their FAQs...
- dragonwriter 3y agoThe point of text to image model is for them to accept natural language (yes, in practice, they all benefit from specialized prompting done with an understanding of model quirks, but that’s not the goal.)
- epivosism 3y agoHere are some play markets on manifold markets tracking its release: https://manifold.markets/markets?s=relevance&f=all&q=deepfloyd https://manifold.markets/markets?s=relevance&f=all&q=deepflo... 35% to full release by end of month, although it may not have adjusted.
- epivosism 3y agoThere's a discord with tons of sample images, where we've been waiting patiently for the release, coming SOON, for 3 months now. https://discord.gg/pxewcvSvNx https://discord.gg/pxewcvSvNx
- CamperBob2 3y agoWhat these AI companies need are some good old-fashioned leakers. We should be seeing these models show up on sketchy pirate sites, complete with garish 80s-style cracking screens crediting various '1337 haX0rs with witty pseudonyms.
- ronsor 3y agoWell, NovelAI was hacked and had their image generation model leaked last year.
- deleted 3y ago[deleted]