27 ms·
Opendream: A layer-based UI for Stable Diffusion
- smrtinsert 3y agoThere's great articles on how layered uis are a lot easier to use than node based uis. Really excited to see a layered approach to SD. Its definitely time to break out of gradio.
- TeMPOraL 3y agoMaybe if they're talking about layered UIs with layer groups, which turn a flat stack into something resembling a tree. But even these UIs don't give you proper non-destructive editing - anything more complex requires you to duplicate parts of layer stack to feed as inputs, which is a destructive operation with respect to structure (those pasted layers won't update if you make changes to copied source). Doing this properly requires a DAG, at which point you're at node-based UIs (or some idiosyncratic mess of an UI that pretends it's not modelling a DAG). It's all moot though, because as far as I know, there is no proper 2D graphics editing software that uses DAGs and nodes. Everyone just copies Photoshop. Especially Affinity, which is grating, given their recent focus on non-destructive editing. For some reason, node-based UIs ended up being a mainstay of VFX, 3D graphics, and VFX & gamedev support tooling. But general 2D graphics - photo editing, raster and vector creation? Nodes are surprisingly absent.
- orbital-decay 3y agoFor some reason, node-based UIs ended up being a mainstay of VFX, 3D graphics, and VFX & gamedev support tooling. But general 2D graphics - photo editing, raster and vector creation? Nodes are surprisingly absent. That's because non-destructive editing is mostly useful for animation, image series/sequences, and asset reuse, which are the most common in these fields. 2D artists have a different mental model, which is additionally set in stone by Photoshop and other software imitating it. Photographers use non-destructive editing, but mostly in simple cases because advanced things (retouching, creative compositing) can't and don't need to be done procedurally anyway.
- danwills 3y agoWhat about ancient 'Illusion', old 'Shake' or current 'Nuke' VFX compositing softwares that totally have been supporting node-based (ie DAG-based) comp-workflows since the early 2000s? Guess this is just a very different (much smaller) realm than your usual Photoshop's and so on?
- dragonwriter 3y ago> There's great articles on how layered uis are a lot easier to use than node based uis I can see that being sensible for simple linear flows from one step to the next, with no branching merging, or connections that skip steps. Seems to me that with any of those other things, a layered UI is going to start to break down a lot faster.
- rytill 3y agoCan you share such articles?
- asynchronous 3y agoVery cool honestly, seems like a much needed improvement over Automatic. Does it support LoRa/will it support in near future?
- varunshenoy 3y agoYou can write an extension to support LoRA (~10 lines of Python HF Diffusers code). If you get to this before me, please create a PR!
- toenail 3y agoFirst thoughts, how do I bind to an ip, and where can I install models?
- gatane 3y agoIs this related to Melondream?
- antman 3y agoCan you add a layer with e.g. an image of yourself?
- ttul 3y agoPretty sure you can do this. Diffusion models by default start with noise, but you can start with any data, including an existing image. For instance, you could import a photo of yourself, mask the eyes and then ask the model to make them green.
- adventured 3y agoNot a bad start. One quick suggestion: avoid the temptation to make it overly complex. Stable Diffusion needs to go out to the masses to a greater degree. The unnecessary garbage complexity (eg Comfy's ridiculous noodlescape) that developers keep including into the UIs is holding Stable Diffusion back significantly from a greater mass adoption.
- bavell 3y agoNode based workflows with little DRY capability (i.e. ComfyUI) do get painful as the workflow grows. That said, an http server capabable of executing ML DAGs is extremely useful and a great building block for other tools and UIs to be built upon. I wrote a typescript API generator for ComfyUI recently and having programmatic access to let you build and send the execution graphs is a game changer. Hoping to have time to release it soon. Same can easily be done for any other language. Exciting stuff!
- tomalaci 3y agoI haven't followed diffusion image generation development for a while. Where do you find information on what models you can use in the model_ckpt field? Do I need to import them from somewhere? What are the main differences between them and which are more modern or better?
- nickstinemates 3y agoYou can find them on huggingface, or you can reverse engineer which ckpt you want to use based on an image you've seen generated (like at majin[1] - beware, there's a lot of NSFW/controversial stuff here.) 1: https://majinai.art/ https://majinai.art/
- bavell 3y agoAlso CivitAI but beware the NSFW https://civitai.com/ https://civitai.com/
- CSSer 3y agoSome of this is straight up soft-core child porn. This is fucked up.
- greggsy 3y agoI believe illustrations have been deemed to be abuse material, so I wouldn’t be surprised if LE have started looking into it.
- CSSer 3y agoGeez that’s disturbing. I clicked having no qualms with nudes, artistic or otherwise. I’m not a prude. I’ve seen my fair share of anime girls and AI nudes. Hell, I was raised on the internet before parental settings were a thing, but I didn’t expect that. It’s so gross how it toes a line too.
- dingnuts 3y ago
- ryukoposting 3y agoIf it can handle LoRAs, I'll be sure to try it out this weekend.
- varunshenoy 3y agoLoRAs can be handled as a straight-forward Python extension!
- _sys49152 3y agoits gonna be breathtaking when this technology gets close enough to make legit cartoons and animations. layers is a step closer to getting there.
- synapticpaint 3y agoThis technology is already close to making animation. Check out some of my experiments with text to video here: https://www.youtube.com/watch?v=CgKNTAjQpkk https://www.youtube.com/watch?v=CgKNTAjQpkk https://youtu.be/X0AhqMhEe-c https://youtu.be/X0AhqMhEe-c
- deleted 3y ago[deleted]
- etra0 3y agoCorridor Crew did a some sort of anime using this technique [1] and then they did two videos [2, 3] explaining the technology behind. Quite interesting if you ask me! There still are some issues with the eyes and a bit of flickering but at the speed everything is moving I wouldn't be surprised if this improves in a year or two. Needless to say, there's still a lot of artistry involved in such a process so anything is yet to be completely automated. [1] https://www.youtube.com/watch?v=tWZOEFvczzA https://www.youtube.com/watch?v=tWZOEFvczzA [2] https://www.youtube.com/watch?v=FQ6z90MuURM https://www.youtube.com/watch?v=FQ6z90MuURM [3] https://www.youtube.com/watch?v=mUFlOynaUyk https://www.youtube.com/watch?v=mUFlOynaUyk
- kitanata 3y ago[flagged]
- renewiltord 3y agoI get that you're spamming out of outrage, but they allowed me to disentangle my comments from my username, which is the same unless you mentioned something you don't want to mention.
- kitanata 3y agoI do not want to disentangle my name from my comments. I want to delete my comments. They are MY comments. I have a right to have them deleted.
- Zuiii 3y agoIf you're in the EU or are willing to stay in the EU for an extended time (>6 months?), then you may be able to compel them to delete the comments using EU laws. If they refuse, escalate and let the entirety of the EU take care of the rest. I get why hn is against deleting comments, and sure, make it really hard to delete comments if necessary, but you should honor the request of users who's unpaid contributions make your site what it is.
- kitanata 3y agoMaybe this is the start of a movement here to get us our own GDPR? You know what’s way more dangerous than me spamming this site with LLM hallucinations? Me walking into Congress with a fucking bill. How far do you want this to go HN? Delete my comments and then delete my account or maybe I start talking to Senators. It’s a nice site you have here. It would a shame to see your data go “poof”, wouldn’t it? You don’t delete my stuff? Maybe I just burn it all down in a nice fire led by Congress and a stroke of a pen. How would you like that? (Also thank you for the support Zuiii. You’re alright with me. :) )
- cateye 3y agoIt makes more sense to embed stable diffusion capabilities into well-established image editors such as Gimp, Photoshop, Krita, or Figma, which come with layered, non-destructive functionalities, rather than attempting the opposite approach. https://github.com/Interpause/auto-sd-paint-ext https://github.com/Interpause/auto-sd-paint-ext https://github.com/thndrbrrr/gimp-stable-boy https://github.com/thndrbrrr/gimp-stable-boy https://www.magicbrushai.com/ https://www.magicbrushai.com/
- j-a-a-p 3y agoDepends on where the 'opposite approach' is aimed to end. If the result is a totally new creative workflow then what is the point of carrying all the ballast of a legacy tool?
- _6atf 3y agoI got briefly very excited for non-destructive editing in GIMP, but the website still says this is slated for 3.2. Which functionality were you referring to?
- 112233 3y agoGimp is so well established that it has almost fossilized... Also, "Normal" layered non-destructive operations are a couple of orders of magnitude faster and do not require 8Gb of VRAM per 512x512 patch, or work only with fixed set of buffer sizes, or any of other strange things SD comes with. Like, how a non-destructive controlnet layer would look in Gimp?
- magic_hamster 3y agoWhile automatic1111 is cumbersome and takes s while to learn, it seems far more capable. The layers here are just inpainting (as noted in the repository readme as well).
- HeartStrings 3y agoHow is this better than A1111?
- deleted 3y ago[deleted]
- denvrede 3y agoThat looks pretty nice but I guess the HW required or the time you have to wait to iterate on these things (if you don't use external services) is quite high. Is there an estimation / idea when a "normal" person can play around with these things with a lot of operational or capital investment?
- cwkoss 3y agoVery cool. Would be interesting to train a model on images with alpha channels so outputs would be automatically masked and more easily composable. But maybe masking is so good these days that would be futile? When a user does img-2-img on a layer does it use the context from other visible layers in the generation?
- dheera 3y agoFor composing this approach works pretty well, maybe the author should consider making a UI for it https://multidiffusion.github.io/ https://multidiffusion.github.io/
- mottiden 3y agoThanks for posting. Really interesting
- Zetobal 3y agoSegmentation is solved... https://github.com/RockeyCoss/Prompt-Segment-Anything https://github.com/RockeyCoss/Prompt-Segment-Anything
- michaelt 3y agoSegment Anything is neat, but segmentation is far from solved. If the user generates a picture of a horse and rider to add onto another composition - they probably want to include the saddle.
- __loam 3y ago[flagged]
- dang 3y agoMaybe so, but please don't post unsubstantive comments to Hacker News.
- __loam 3y ago[flagged]
- Zuiii 3y agoIt's comments like this that really highlight the extent of damage copyright has caused. It has conditioned people to think that they can own information once released, and that they can treat it like property. It's a notion that's ridiculous and silly that it gets threatened each time we make a leap in technology (tapes, computers, internet and streaming, and now generative AI). I wonder how long society will tolerate the nonsensical idea of copyright before it's had enough. Intellectual property is nonsense and I'm glad it keeps getting exposed.
- __loam 3y agoUntil we have an economy where people don't starve to death from lack of food or die of frostbite from being homeless, we have to figure out ways for people who actually make things to benefit from that work. Copyright is currently protecting people who rely on the product of their labor to make ends meet. Until we have the fairy tale economy you envision where artists can have the work stolen and still live, this is the best we've got. And yeah, wild to think people feel entitled to own their own work thanks to silly things like the entire body of copyright law.
- Zuiii 3y agoThe true fairy tale is the intellectual "property" nonsense you people deluded yourself into believing. As I said, reality will continue to slap you across the face with each jump in technology. How I wish these slaps were enough to wake you up from your delusions but alas.
- smallerfish 3y agoSlap a virtualenv setup into that install script please. A system wide pip install is a bad pattern.
- deleted 3y ago[deleted]
- varunshenoy 3y agodone :)
- noman-land 3y agoNow that's agile.
- Hamcha 3y agoWhat's up with names nowadays? Not only there's already an OpenDream[1] on GitHub, but there's also a Stable Diffusion service also called OpenDream[2]! 1. https://github.com/OpenDreamProject/OpenDream https://github.com/OpenDreamProject/OpenDream 2. https://opendream.ai/ https://opendream.ai/
- deleted 3y ago[deleted]
- tavavex 3y agoVery exciting. The "first-generation" Stable Diffusion frontends seem to have settled on a specific design philosophy, so it's interesting to see new tools (like this or ComfyUI) shake up the way people work with this tool. I hope that in a few years, we'll know which philosophy works best.
- TillE 3y agoOut of all the AI-related tools, generative art frontends are probably the thing most likely to radically change and improve in the next few years. It's specifically why I've avoided diving too deep into "prompt engineering", because the kind of incantations required today just aren't going to be the way most people interact with this stuff for very long.
- bobboies 3y agoIncantations are fun!
- greggsy 3y agoIt’s entirely likely that there’s much more effort going into generative text - any perceived advancement of generative images is going to be disproportionately skewed due the richness of information that they hold.
- orbital-decay 3y ago> Out of all the AI-related tools, generative art frontends are probably the thing most likely to radically change and improve in the next few years. The difference between UIs is actually not very relevant today; by now the generic workflow for complex scenes is more or less obvious to anyone who spent time with SD. - Draw basic composition guides. Use them with controlnets or any other generic guidance method to enforce the environment composition you want. Train your own controlnet if you need something specific. (lots of untapped potential here) - Finetune the checkpoint on your reference pictures or use other style transfer methods to enforce the consistent style. - Use manual brush masking, manually guided segmentation (ex. SAM), or prompted segmentation (ex ClipSEG) to select the parts to be replaced with other objects. The choice depends on your case and need to do it procedurally. - Photobash and add detail to the elements of your scene using any composition methods you have (noisy latent composition, inpainting etc) with the masks you created in the previous step. Use advanced guidance (controlnets, t2i adapters etc) - Don't bother with any prompts beyond very basic descriptions, as "prompt engineering" is slow and unreliable. Don't overwhelm the model by trying to fit lots of detail in one pass; use separate passes for separate objects or regions. - Alternative 3D version: build a primitive 3D scene from basic props (shapes, rigs). Render the backdrop and separate objects into separate layers as guides. Use them with controlnets & co to render the scene in a guided manner, combining the objects by latent composition, inpainting, or any other means. This can be used for procedural scenes and animation (although current models lack temporal stability). As long as your tool has all that in one place, it's a breeze, regardless of the UI paradigm (admittedly auto1111's overloaded gradio looks straight out of a trash compactor nowadays). I expect 2D/3D software integrations being the most successful in the future, as they already offer proven UIs and most desirable side features. The problem is that in the current state SD can't do much in the production setting, it's not a finished product - so there's not a lot of interest in software integrations just yet.
- brianjking 3y agoIs it possible to add SD XL support for this? I'd love a colab notebook if anyone has the skill and time to do so.
- varunshenoy 3y agoIf anyone wants to add SDXL support, all you have to do is create a new extension with the correct SDXL logic (loading from HF diffusers, etc.). You could parameterize `num_inference_steps`, for example, to delegate decisions to the user of the extension. If anyone gets to making one before me, please leave a PR!