58 ms·
Stable Diffusion 3
- poulpy123 3y agoDidn't they released another model few days ago ?
- satisfice 3y agoCan it make a picture of a woman chasing a bear? The old one can't.
- ShamelessC 3y ago[flagged]
- jetrink 3y agoIt is a challenge for these models to generate images of counterintuitive or unusual situations that aren't depicted in the training set. For example, if you ask for a small cube sitting on top of a large cube, you'll likely get the correct result on the first attempt. Ask for a large cube on a small cube and you'll probably get an image of them side-by-side or with the small cube on top instead. The models can generalize in impressive ways, but it's still limited.
- edflsafoiewq 3y agoIs it a matter of being depicted in the training data, or is it a matter of understanding the grammatical relationships in the prompt?
- Majromax 3y agoIt's likely a result of the interplay between the image generation and caption/description generation aspects of the model. The earliest diffusion-based image generators used a 'bag of words' model for the caption (see musing regarding this and DALL-E 3: https://old.reddit.com/r/slatestarcodex/comments/16y14co/scott_has_won_his_ai_image_bet/k36psm7/ https://old.reddit.com/r/slatestarcodex/comments/16y14co/sco...), whereby 'a woman chasing a bear' would turn into `['a', 'a', 'chasing', 'bear', 'woman']`. That's good enough to describe compositions well-represented in the training set, but it will be likely to lock-in to those common representations at the expense of rarer but still possible ones (the 'woman chasing a bear' above).
- isaacfrond 3y agoTried it myself. Couldn't do it. It's a nice test case!
- obloid 3y agoA while ago my daughter wanted an image of Santa pulling a sleigh with a reindeer in the driver's seat holding the reins. We tried dozens of different prompts and Dall-e 3 could not do it.
- brandall10 3y agoBeing able to generate content w/ minimal presence in the training set is arguably an emergent, desirable behavior that could be seen as a form of intelligence.
- cheald 3y agoSD 1.5 (using RealisticVision 5.1, 20 steps, Euler A) spit out something technically correct (but hilarious) in just a few generations. "a woman chasing a bear, pursuit" https://i.imgur.com/RqCXVYC.png https://i.imgur.com/RqCXVYC.png
- pqdbr 3y agoThe sample images are absolutely stunning. Also, I was blown away by the "Stable Diffusion" written on the side of the bus.
- kzrdude 3y agoIs it just me or is the stable diffusion bus image broken in the background? The bus back there does not look logical w.r.t placement and size relative to the sidewalk.
- PcChip 3y agoThe text/spelling part is a huge step forward
- gat1 3y agoI guess we do not know anything about the training dataset ?
- _1 3y agoIt's ethical
- kranke155 3y ago"Ethical"
- amirhirsch 3y agoThe dataset is so ethical that it is actually just a press release and not generally available.
- wtcactus 3y agoWho decides what's ethical in this scenario? Is it some independent entity?
- potwinkle 3y agoI decided.
- deleted 3y ago[deleted]
- thelazyone 3y agoThis is a good question - not only for the actual ethics of the training, but for the future of AI use for art. It's both gonna damage the livelyhood of many artists (me included, probably) but also make it accessibly to many more people. As long as the training dataset is ethical, I think fighting it is hard and pointless.
- yreg 3y agoWhat data would you consider making the dataset unethical vs. ethical?
- kbumsik 3y agoSo there is no license information yet?
- alexb_ 3y ago> We believe in safe, responsible AI practices. This means we have taken and continue to take reasonable steps to prevent the misuse of Stable Diffusion 3 by bad actors. Safety starts when we begin training our model and continues throughout the testing, evaluation, and deployment. In preparation for this early preview, we’ve introduced numerous safeguards. By continually collaborating with researchers, experts, and our community, we expect to innovate further with integrity as we approach the model’s public release. What exactly does this mean? Will we be able to see all of the "safeguards" and access all of the technology's power without someone else's restrictions on them?
- Tiberium 3y agoFor SDXL this meant that there were almost no NSFW (porn and similar) images included in the dataset, so the community had to fine-tune the model themselves to make it generate those.
- hhjinks 3y agoThe community would've had to do that anyway. The SD1.5-based NSFW models of today are miles ahead of those from just a year ago.
- Der_Einzige 3y agoAnd the pony SDXL nsfw model is miles ahead of SD1.5 NSFW models. Thank you bronies!
- tifi 3y ago[flagged]
- sschueller 3y agoNo worries, the safeguards are only for the general public. Criminals will have no issues going around them. /s
- 3y ago
- willsmith72 3y agoat this point perfect text would be a gamechanger if it can be solved midjourney 6 can be completely photorealistic and include valid text, but also sometimes adds bad text. it's not much, but having to use an image editor for that is still annoying. for creating marketing material, getting perfect text every time and never getting bad text would be amazing
- falcor84 3y agoI wonder if we could get it to generate a layered output, to make it easy to change just the text layer. It already creates the textual part in a separate pass, right?
- deprecative 3y agoI would bet that Adobe is definitely salivating at that. Might not be for a long time but it seems like a no brainer once the technology can handle it. Just the last few years have been fast and I interacted with the JS landscape for a few years. It moves faster than Sonic and this tech iterates quick.
- spywaregorilla 3y agoCurrent open source tools include pretty decent off the shelf segment anything based detectors. It leaves a lot to be desired, but you do layer-like operations automatically detecting certain concept and applying changes to them or, less commonly exporting the cropped areas. But not the content "beneath" the layers as they don't exist.
- snovv_crash 3y agoWhich tools would you recommend for this kind of thing?
- spywaregorilla 3y agocomfyui + https://github.com/ltdrdata/ComfyUI-Impact-Pack https://github.com/ltdrdata/ComfyUI-Impact-Pack
- FloatArtifact 3y agoI'm curious to know if they're safeguards are eliminated when users find tune the model?
- patates 3y agoHalf of the announcement talks about safety. The next step will be these control mechanisms being built into all sorts of software I suppose. It's "safe" for them, not for the users, at least they should make that clear.
- spir 3y agothanks, i hadn't fully realized that 'safety' means 'safe to offer' and not 'safe for users'. i won't forget it
- wiz21c 3y agoThey rather talk about "reasonable steps" to safety. Sounds like "just the minimum so we don't end up in legal trouble" to me...
- tasty_freeze 3y agoThere is some truth in what you say, just like saying you're a "free speech absolutist" sounds good at first blush. But the real world is more complicated, and the provider adds safety features because they have to operate in the real world and not just make superficial arguments about how things should work. Yes, they are protecting themselves from lawsuits, but they are also protecting other people. Preventing people asking for specific celebrities (or children) having sex is for their benefit too.
- s1k3s 3y agoI truly wonder what "unsafe" scenarios an image generator could be used for? Don't we already have software that can do pretty much anything if a professional human is using it?
- t_von_doom 3y agoI would say the barrier to entry is stopping a lot of ‘candid’ unsafe behaviour. I think you allude to it yourself in implying currently it requires a professional to achieve the same results. But giving that ability to _everyone_ will lead to a huge increase in undesirable and targeted/local behaviour. Presumably it enables any creep to generate what they want by virtue of being able to imagine it and type it, rather than learn a niche skill set or employ someone to do it (who is then also complicit in the act)
- 123yawaworht456 3y ago>This preview phase, as with previous models, is crucial for gathering insights to improve its performance and safety ahead of an open release. oh, for fuck's sake.
- memossy 3y agoWe did this for every stable diffusion release, you get the feedback data to improve it continuously ahead of open release.
- 123yawaworht456 3y agoI was referring to 'safety'. how the hell can an image generation model be dangerous? we had software for editing text, images, videos and audio for half a century now.
- Jensson 3y agoAdvertisers will cancel you if you do anything they don't like, 'safety' is to prevent that.
- glimshe 3y agoThis reinforces my impression that Google is at least one year behind. Stunning images, 3D, video while Gemini had to be partially halted this morning.
- bamboozled 3y agoFor "political" reasons, not for technical reasons. Don't get it twisted.
- coeneedell 3y agoI would describe those issues as technical. It’s genuinely getting things wrong because the “safety” element was implemented poorly.
- anononaut 3y agoThose are safety elements which exist for political reasons, not technical ones.
- ethbr1 3y agoOf all criticism that could be leveled at Google, 'shipping a product and supporting it' being the only thing that matters seems fair. Which takes all the behind the scenes steps, not just the technical ones.
- deleted 3y ago[deleted]
- verticalscaler 3y agoYou think that technology is first. You think that mathematicians and computer engineers or mechanical engineers or doctors are first. They’re very important, but they’re not first. They’re second. Now I’ll prove it to you. There was a country that had the best mathematicians, the best physicists, the best metallurgists in the world. But that country was very poor. It’s called the Soviet Union. But when you took one of these mathematicians or physicists, who was smuggled out or escaped, put him on a plane and brought him to Palo Alto. Within two weeks, they were producing added value that could produce great wealth. What comes first is markets. If you have great technology without markets, without a market-friendly economy, you’ll get nowhere. But if you have a market-friendly economy, sooner or later the market forces will give you the technology you want. And that my friend, simply won't come from an office paralyzed by internal politics of fear and conformity. Don't get it twisted.
- treesciencebot 3y agoQuite nice to see diffusion transformers [0] becoming the next dominant architecture on the generative media. [0]: https://twitter.com/EMostaque/status/1760660709308846135 https://twitter.com/EMostaque/status/1760660709308846135
- amelius 3y agoDoes anyone know of a good tutorial on how diffusion models work?
- Ologn 3y agoI liked this 18 minute video ( https://www.youtube.com/watch?v=1CIpzeNxIhU https://www.youtube.com/watch?v=1CIpzeNxIhU ). Computerphile has other good videos with people like Brian Kernighan.
- spaceheater 3y agofast.ai has a whole free course https://www.youtube.com/watch?v=_7rMfsA24Ls https://www.youtube.com/watch?v=_7rMfsA24Ls https://course.fast.ai/Lessons/part2.html https://course.fast.ai/Lessons/part2.html
- jasonjmcghee 3y agohttps://jalammar.github.io/illustrated-stable-diffusion/ https://jalammar.github.io/illustrated-stable-diffusion/ His whole blog is fantastic. If you want more background (e.g. how transformers work) he's got all the posts you need
- amelius 3y agoThis looks nice, thank you, but I'm looking for a more hands-on tutorial, with e.g. Python code, like Andrej Karpathy makes them.
- keiferski 3y agoThe obsession with safety in this announcement feels like a missed marketing opportunity, considering the recent Gemini debacle. Isn’t SD’s primary use case the fact that you can install it on your own computer and make what you want to make?
- jsheard 3y agoAt some point they have to actually make money, and I don't see how continuously releasing the fruits of their expensive training for people to run locally on their own computer (or a competing cloud service) for free is going to get them there. They're not running a charity, the walls will have to go up eventually. Likewise with Mistral, you don't get half a billion in funding and a two billion valuation on the assumption that you'll keep giving the product away for free forever.
- keiferski 3y agoBut there are plenty of other business models available for open source projects. I use Midjourney a lot and (based on the images in the article) it’s leaps and bounds beyond SD. Not sure why I would switch if they are both locked down.
- Auracle 3y agoSD would probably be a lot better if they didn't have to make sure it worked on consumer GPUs. Maybe this announcement is a step towards that where the best model will only be able to be accessed by most using a paid service.
- raxxorraxor 3y agoI believe the opposite. I think the ability for people to adopt models made SD more successful than any other model for image synthesis in the first place. Similarly how consumer PCs drove innovation towards faster hardware. I believe it to be the reference of image synthesis for that matter, so "better" is a bit blurry.
- lreeves 3y agoPeople in this discussion seem to be hand-wringing about Stability's "saftey" comments but every model they've released has been fine tuned for porn in like 24 hours.
- mopierotti 3y agoThat's not entirely true. This wasn't the case for SD 2.0/2.1, and I don't think SD 3.0 will be available publicly for fine tuning.
- lreeves 3y agoSD 2 definitely seems like an anomaly that they've learned from though and was hard for everyone to use for various reasons. SDXL and even Cascade (the new side-project model) seems to be embraced by horny people.
- viraptor 3y ago2 is not popular because people have better quality results with 1.5 and xl. That's it. If 3 is released and works better, it will be fine tuned too.
- londons_explore 3y agoAll the demo images are 'artwork'. will the model also be able to produce good photographs, technical drawings, and other graphical media?
- spywaregorilla 3y agoPhotorealism is well within current capabilities. Technical drawings absolutely not. Not sure what other graphical media includes.
- sweezyjeezy 3y agoYeah but try getting e.g. Dall-E 3 to do photorealism, I think they've RLHF'd the crap out of it in the name of safety.
- spywaregorilla 3y agowell that's what you get with closed ai.
- DiscourseFan 3y agoThat's why we need open AI which scoops up all the data with its specific contexts and history and transforms it into a vast incomprehensible machine for us peons to gawk at while we starve and boil to death
- spywaregorilla 3y agolow quality discourse imo
- astrange 3y agoThat's not safety, the safety RLHF is because it tries to generate porn and people with three legs if you don't stop it. It has the weird art style because that's what looks the most "aesthetic". And because it doesn't actually have nearly as good enough data as you'd think it does. Sora looks like it could be better.
- londons_explore 3y agoI really wonder what harm would come to the company if they didn't talk about safety? Would investors stop giving them money? Would users sue that they now had PTSD after looking at all the 'unsafe' outputs? Would regulators step in and make laws banning this 'unsafe' AI? What is it specifically that company management is worried about?
- brainwipe 3y agoAll of the above! Additionally... I think AI companies are trying to steer the conversation about safety so that when regulations do come in (and they will) that the legal culpability is with the user of the model, not the trainer of it. The business model doesn't work if you're liable for harm caused by your training process - especially if the harm is already covered by existing laws. One example of that would be if your model was being used to spot criminals in video footage and it turns out that the bias of the model picks one socioeconomic group over another. Most western nations have laws protecting the public against that kind of abuse (albeit they're not applied fairly) and the fines are pretty steep.
- graphe 3y agoThey have already used "AI" with success to give people loans and they were biased. Nothing happened legally to that company.
- dorkwood 3y agoThey're attempting to guard themselves against incoming regulation. The big players, such as Microsoft, want to squash Stable Diffusion while protecting themselves, and they're going to do it by wielding the "safety is important and only we have the resources to implement it" hammer.
- HeatrayEnjoyer 3y agoSafety is a very real concern, always has been in ML research. I'm tired of this trite "they want a moat" narrative. I'm glad tech orgs are for once thinking about what they're building before putting out society-warping democracy-corroding technology instead of move fast break things.
- inference-lord 3y agoCool but it's hard to keep getting "blown away" at this stage. The "incredible" is routine now.
- danparsonson 3y agoSo... they should just stop?
- inference-lord 3y agoDidn't say that, I just mean, it's great to have a new demo each week, but it's getting less and less exciting. Same things over and over again.
- dougmwne 3y agoAt this point, the next thing that will blow me away is AGI at human expert level or a Gaussian Splat diffusion model that can build any arbitrary 3D scene from text or a single image. High bar, but the technology world is already full of dark magic.
- attilakun 3y agoIs there a Guassian splat model that works without the "Structure from Motion" step to extract the point cloud? That feels a bit unsatisfying to me.
- inference-lord 3y agoWill ask it for immortality, endless wealth, and still get bored.
- consumer451 3y agoI would be a big fan of solid infographics or presentation slides. That would be very useful.
- JonathanFly 3y agoFrom: https://twitter.com/EMostaque/status/1760660709308846135 https://twitter.com/EMostaque/status/1760660709308846135 Some notes: - This uses a new type of diffusion transformer (similar to Sora) combined with flow matching and other improvements. - This takes advantage of transformer improvements & can not only scale further but accept multimodal inputs.. - Will be released open, the preview is to improve its quality & safety just like og stable diffusion - It will launch with full ecosystem of tools - It's a new base taking advantage of latest hardware & comes in all sizes - Enables video, 3D & more.. - Need moar GPUs.. - More technical details soon >Can we create videos similar like sora Given enough GPUs and good data yes. >How does it perform on 3090, 4090 or less? Are us mere mortals gonna be able to have fun with it ? Its in sizes from 800m to 8b parameters now, will be all sizes for all sorts of edge to giant GPU deployment. (adding some later replies) >awesome. I assume these aren't heavily cherry picked seeds? No this is all one generation. With DPO, refinement, further improvement should get better. >Do you have any solves coming for driving coherency and consistency across image generations? For example, putting the same dog in another scene? yeah see @Scenario_gg's great work with IP adapters for example. Our team builds ComfyUI so you can expect some really great stuff around this... >Dall-e often doesn’t even understand negation, let alone complex spatial relations in combination with color assignments to objects. Imagine the new version will. DALLE and MJ are also pipelines, you can pretty much do anything accurately with pipelines now. >Nice. Is it an open-source / open-parameters / open-data model? Like prior SD models it will be open source/parameters after the feedback and improvement phase. We are open data for our LMs but not other modalities. >Cool!!! What do you mean by good data? Can it directly output videos? If we trained it on video yes, it is very much like the arch of sora.
- cheald 3y agoSD 1.5 is 983m parameters, SDXL is 3.5b, for reference. Very interesting. I've been streching my 12GB 3060 as far as I can; it's exciting that smaller hardware is still usable even with modern improvements.
- memossy 3y ago800m is good for mobile, 8b for graphics cards. Bigger than that is also possible, not saturated yet but need more GPUs.
- coldcode 3y agoNo details in the announcement, is it still pixel size in = pixel size out?
- spywaregorilla 3y agoImpressive text in the images.
- btbuildem 3y agoThat's nice, but could we please have an unsafe alternative? I would like to footgun both my legs off, thank you.
- dougmwne 3y agoSince these are open models, people can fine tune them to do anything.
- politician 3y agoIt’s not obvious that fine-tuning can remove all latent compulsions from these models. Consider that the creators know that fine-tuning exists and have vastly more resources to explore the feasibility of removing deep bias using this method.
- dougmwne 3y agoGo check out the Unstable Diffusion Discord.
- SV_BubbleTime 3y agoThe vast majority of images there are SD1.5, even the ones made today. Which goes far more towards the idea that safety isn’t a desirable feature to a lot of AI users.
- ttul 3y agoI suppose you could train a model from scratch if you have enough money to blow…
- Sharlin 3y ago[flagged]
- btbuildem 3y agoConsider the impact of the two scenarios below: 1. A mass-distributed LLM (hosted by google or openai or whoever) that's been neutered and twisted into political correctness in a haphazard series of kneejerk meetings of small groups of people who are terrified big investors will walk away or that some powerful political group will denounce them. Effectively they create an enormous bias of falsity and incorrectness, for billons of people to use and embed the results throughout all their intellectual output. 2. Some wacko with an expensive Nvidia GPU makes deepfake porn of a popular politician. Or goodness forbid, of some weird kink where if this was actually a scene filmed with real people, there would be serious ethical issues. Which scenario do you think is more dangerous, long term, and in terms of broad impact on society in general?
- deepsdev 3y agoCan we use it create SORA like videos?
- memossy 3y agoIf we trained it with videos yes but need more GPUs for that.
- nickthegreek 3y agoNo.
- wtcactus 3y agoI notice they are avoiding images of people in the announcement. I wonder if they are afraid of the same debacle as google AI and what they mean by "safety" is actually heavy bias against white people and their culture like what happened with Gemini.
- danielbln 3y agoWhat's white people culture?
- potwinkle 3y agoFrom the examples I see on Twitter, they are usually referring to the different cultures of Irish, European, and American white people. Gemini, in an effort to reverse the bias that the models would naturally have, ends up replacing these people with those from other cultures.
- astrange 3y agoCalling Irish people white is a rather historically radical statement.
- sealeck 3y agoWhite is a pretty complex and non-obvious category.
- int_19h 3y agoSince the definition of "white" is inherently cultural, it varies from place to place and from time to time. Today, in US and Europe, pretty much everyone who cares about racial categorization would consider Irish "white". Historically, it was different, but that is only relevant when discussing history.
- samatman 3y agoNot true, that was made up by race baiting academics who were lying liars. It's absurd that anyone fell for it, frankly. The lie originates with a Communist race hustler named Noel Ignatiev, also known for publishing Race Traitor magazine. A thoroughly unpleasant person.
- cuckatoo 3y agoNSFW fine tune when? Or will "safety" win this time?
- SXX 3y agoThey need to release model first. Then it's will be fine-tuned.
- redder23 3y agoHorrible website, hijacks scrolling. I have my scrolling speed up with Chromium Wheel Smooth Scroller. This website's scrolling is extremely slow, so the extension is not working because they are "doing it wrong" TM and somehow hijack native scrolling and do something with it.
- pama 3y agoI wish they put out the report already. Has anyone else published a preprint combining ideas similar to diffusion transformers and flow matching?
- redder23 3y ago[flagged]
- Tadpole9181 3y agoThey are a business and operate in the objective reality that a product that can generate imagery like child porn draws intense scrutiny from investors and law makers. Not to mention they have their own moral compass that they don't feel comfortable giving away a tool that they feel can do harm. You're free to start your own competing business.
- subzel0 3y ago“Photo of a red sphere on top of a blue cube. Behind them is a green triangle, on the right is a dog, on the left is a cat” https://pbs.twimg.com/media/GG8mm5va4AA_5PJ?format=jpg&name=large https://pbs.twimg.com/media/GG8mm5va4AA_5PJ?format=jpg&name=...
- Workaccount2 3y agoNot bad, I'm curious of the output if you ask for a mirrored sphere instead.
- svenmakes 3y agoThis is actually the approach of one paper to estimate lighting conditions. Their strategy is to paint a mirrored sphere onto an existing image: https://diffusionlight.github.io/ https://diffusionlight.github.io/
- jetrink 3y agoOne thing that jumps out to me is that the white fur on the animals has a strong green tint due to the reflected light from the green surfaces. I wonder if the model learned this effect from behind the scenes photos of green screen film sets.
- diggan 3y agoIt's just diffuse irradiance, visible in most real (and CGI) pictures although not as obvious as that example. Seems like a typical demo scene for a 3D renderer, so I bet that's why it's so prominent.
- zero_iq 3y agoThe models do a pretty good job at rendering plausible global illumination, radiosity, reflections, caustics, etc. in a whole bunch of scenarios. It's not necessarily physically accurate (usually not in fact), but usually good enough to trick the human brain unless you start paying very close attention to details, angles, etc. This fascinated me when SD was first released, so I tested a whole bunch of scenarios. While it's quite easy to find situations that don't provide accurate results and produce all manner of glitches (some of which you can use to detect some SD-produced images), the results are nearly always convincing at a quick glance.
- the_duke 3y agoSo, they just announced StableCascade. Wouldn't this v3 supersede the StableCascade work? Did they announce it because a team had been working on it and they wanted to push it out to not just lose it as an internal project, or are there architectural differences that make both worthwile?
- Kubuxu 3y agoI think of the SD3 as a further evolution of SD1.5/2/XL and StableCascade as a branching path. It is unclear which will be better in the long term, so why not cover both directions if they have the resources to do so?
- ttul 3y agoI suspect Stable Cascade may incorporate a DiT at some point. The UNet is easily swapped out. SC’s main innovation is the training of a semantic compressor model and a VQGAN that translates the latent output from the diffusion model back to image space - rather than relying on a VAE. It’s a really smart architecture and I think is fertile ground for stacking on new things like DiT.
- whywhywhywhy 3y agoThere's architectural differences, although I found Stable Cascade a bit underwhelming, while yes it can actual manage text, the text it does manage just looks like someone just wrote text over the image it doesn't feel integrated a lot of the time. SD3 seems to be more towards SOTA, not sure why Cascade took so long to get out, seemed to be up and running months ago
- Dwedit 3y agoStable Cascade has a distinct noisy look to generated images. It almost looks as bad as images being dithered to the old 216 color Netscape palette.
- ttul 3y agoIf you renoise the output of the first diffusion stage to halfway and then denoise forward again, you can eliminate the bad output. This approach is called “replay” or “iterative mixing” and there are a few open source nodes for ComfyUI you can refer to.
- 101008 3y agoWhat's the best way to use SD (3 or 2) online? I can't run it on my PC and I want to do some experiments to generate assets for a POC videogame I'm working on. I pay MidJOurney and I woulnd't mind pay something like 5 or 10 dollars per month to experiment with SD, but I can't find anything.
- Gracana 3y agoI used Rundiffusion for a while before I bought a 4090, and I thought their service was pretty nice. You pay for time on a system of whatever size you choose, with whatever tool/interface you select. I think it's worth tossing a few bucks into it to try it out.
- heroprotagonist 3y agoEh, you can get the same software up and running in less than 15-20 minutes on an EC2 GPU instance for about half the hourly-rated pricing of rundiffusion. And you'll also pay less than their 'premium' monthly fee for storage of keeping an instance in the Stopped state the entire month. I used rundiffusion to play around with a bunch of different open source software quickly and easily with pre-downloaded models after getting annoyed at my laptop GPU. But once I settled on one particular implementation and started spending a lot of time in it, it no longer made sense to repeatedly pay every hour for an initial ease-of-setup. The only real ongoing benefit was rundiffusion came with a bunch of models pre-downloaded so swapping between them was quick. But you can use UI addons like the CivitAI browser to download models automatically through automatic1111, and you'll likely want to go beyond what they predownload to the instance for you anyway. The downside to running on the cloud directly is having to manage the running/stopped state of the instance yourself. I haven't ever left it running when I was done with an instance, but I could see that as a risk. CLI commands and scripting can make that faster than logging into a website which does it for you automatically, but it's extra effort. I thought about building an AMI and putting it up on AWS marketplace, but it looks like there are a few options for that already. I don't know how good they are out of the box, as I haven't used them. But if spending 20 minutes once to get software running on a Linux instance is truly the only barrier to reducing cost, those prebuilt AMIs are a decent intermediary step. They're about $0.10/hour on top of server costs. I skipped straight to installing the software myself, but even an extra $0.10/hour overhead would be better than paying double..
- bsaul 3y agoAnyone knows which AI could be used to generate UI design elements ? (such as "generate a real estate app widget list") as well as the kind of prompts one would use to obtain good results ? I'm only now investigating using AI to increase velocity in my projects, and the field is moving so fast, i'm a bit outdated.
- kevinbluer 3y agov0 by Vercel could be worth a look: https://v0.dev https://v0.dev From the FAQ: "v0 is a generative user interface system by Vercel powered by AI. It generates copy-and-paste friendly React code based on shadcn/ui and Tailwind CSS that people can use in their projects"
- gwern 3y agoIf by design elements you include vector images, you could try https://www.recraft.ai/ https://www.recraft.ai/ or Adobe Firefly 2 - there's not a lot of vector work right now, so your choices are either the handful of vector generators, or just bite the bullet and use eg DALL-E 3 to generate raster images you convert to SVG/recreate by hand. (The second is what we did for https://gwern.net/dropcap https://gwern.net/dropcap because the PNG->SVG filesizes & quality were just barely acceptable for our web pages.)
- Auracle 3y agoIt's really unfortunate that Silicon Valley ended up in an area that's so far left - and to be clear, it'd be just as bad if it was in a far right area too. Purple would have been nice, to keep people in check. 'Safety' seems to be actively making AI advances worse.
- asadotzler 3y agoSo far left the techies dont even have a labor union. You're a joke.
- spencerflem 3y agoSilicon Valley is not "far left" by any stretch, which implies socialism, redistribution of wealth, etc. This is obvious by inspection. I assume by far left, you mean progressive on social issues, which is not really a leftist thing but the groups are related enough that I'll give you a pass. Silicon valley techies are also not socially progressive. Read this thread or anything published by Paul Graham or any of the AI leaders for proof of that. However most normal city people are. A large enough percent of the country that big companies that want to make money feel the need to appeal to them. Funnily enough, what is a uniquely Silicon Valley political opinion is valuing the progress of AI over everything else
- TulliusCicero 3y agoTechies are socially progressive as a whole. Yes there are some outliers, and tech leaders probably aren't as far left socially as the ground level workers.
- spencerflem 3y agoI wish :/, I really do I find them in general to not be Republican and all the baggage that entails but the typical techie I meet is less concerned with social issues than the typical city Democrat. If I can speculate wildly, I think it is because tech has this veneer of being an alternative solution to the worlds problems, so a lot of techies believe that advancing of tech is both the most important goal and also politically neutral. And also, now that tech is a uniquely profitable career, the types of people that would be in business majors are now CS majors. Ie. those that are mainly interested in getting as much money as possible for themselves.
- 4bpp 3y agoI guess we should count our blessings and be grateful that literacy, the printing press, computers and the internet became normalised before this notion of "harm" and harm prevention was. Going forward, it's hard to imagine how any new technology that is unconditionally intellectually empowering to the individual will be tolerated; after all, just think of the harms someone thus empowered could be enabled to perpetrate. Perhaps eventually, once every forum has been assigned a trust-and-safety team and word processor has been aligned and most normal people have no need for communication outside the Metaverse (TM) in their daily lives, we will also come around to reviewing the necessity of teaching kids to write, considering the epidemic of hateful graffiti and children being caught with handwritten sexualised depictions of their classmates.
- xanderlewis 3y ago> unconditionally intellectually empowering What makes you think those who’ve worked hard over a lifetime to provide (with no compensation) the vast amounts of data required for these — inferior by every metric other than quantity — stochastic approximations of human thought should feel empowered? I think the genAI / printing press analogy is wearing rather thin now.
- graphe 3y agoWHO exactly worked hard over a lifetime with no compensation?
- xanderlewis 3y agoBy compensation I mean from the companies creating the models, like OpenAI.
- graphe 3y agoComputers and drafters had their work taken by machines. IBM did not pay off the computers and drafters. In this case you could make a steady decent wage. My grandfather was trained in a classic drawing style (yes it was his main job). He did not get into the profession to make money. He did it out of passion and died poor. Artists are not being tricked by the promise of wealth. You will get a cloned style if you can't afford the real artist making it and if the commission goes to a computer how is that not the same as plagerism by a human? Artists were not being paid well before. The anime industry has proven the endpoint of what happens to artists as a profession despite their skills. Chess still exists despite better play by machines. Art as a commercial medium has always been tainted by outside influences such as government, religion and pedophilia. In the end, drawing wasn't going to survive in the age of vector art and computers. They are mainly forgettable jpgs you scroll past in a vast array like DeviantArt.
- hizanberg 3y agoIMO the "safety" in Stable Diffusion is becoming more overzealous where most of my images are coming back blurred, where I no longer want to waste my time writing a prompt only for it to return mostly blurred images. Prompts that worked in previous versions like portraits are coming back mostly blurred in SDXL. If this next version is just as bad, I'm going to stop using Stability APIs. Are there any other text-to-image services that offer similar value and quality to Stable Diffusion without the overzealous blurring? Edit: Example prompt's like "Matte portrait of Yennefer" return 8/9 blurred images [1] [1] https://imgur.com/a/nIx8GBR https://imgur.com/a/nIx8GBR
- nickthegreek 3y agoRun it locally.
- lolinder 3y agoI haven't tried SD3, but my local SD2 regularly has this pattern where while the image is developing it looks like it's coming along fine and then suddenly in the last few rounds it introduces weird artifacts to mask faces. Running locally doesn't get around censorship that's baked into the model. I tend to lean towards SD1.5 for this reason—I'd rather put in the effort to get a good result out of the lesser model than fight with a black box censorship algorithm. EDIT: See the replies below. I might just have been holding it wrong.
- yreg 3y agoDo you use the proper refiner model?
- lolinder 3y agoProbably not, since I have no idea what you're talking about. I've just been using the models that InvokeAI (2.3, I only just now saw there's a 3.0) downloads for me [0]. The SD1.5 one is as good as ever, but the SD2 model introduces artifacts on (many, but not all) faces and copyrighted characters. EDIT: based on the other reply, I think I understand what you're suggesting, and I'll definitely take a look next time I run it. [0] https://github.com/invoke-ai/InvokeAI https://github.com/invoke-ai/InvokeAI
- ametrau 3y ago“Safety” = safe to our reputation. It’s insulting how they imply safety from “harm”.
- kingkawn 3y agoSo they should dash their company on the rocks of your empty moral positions about freedom?
- dingnuts 3y agoshould pens be banned because a talented artist could draw a photorealistic image of something nasty happening to someone real?
- mrighele 3y agoPhotoshop and the likes (modern day's pens) should have an automatic check that you are not drawing porn, censor the image and report you to the authorities if it thinks it involves minors. edit: yes it is sarcasm, though I fear somebody will think it is in fact the right way to go.
- mtlmtlmtlmtl 3y agoThat's ridiculous. What about real pens and paintbrushes? Should they be mandated to have a camera that analyses everything you draw/write just to be "safe"? Maybe we should make it illegal to draw or write anything without submitting it to the state for "safety" analysis.
- gambiting 3y agoI hope that's sarcasm.
- IMTDb 3y agoText editors and the likes (modern day's typewriters) should have an automatic check that you are not criticizing the government, censor the text and report you to the authorities if it thinks it an alternate political party. Hopefully you are going to be absolutely shocked by the prospect of the above sentence. But as you can see, surveillance is a slippery slope. "Safety" is a very dangerous word because everybody wants to be "safe" but no one is really ready to define what "safe" actually means. The moment we start baking cultural / political / environmental preferences and biases in the tools we use to produce content, we allow other group of people with different views to use those "safeguards" to harm us or influence us in ways we might not necessarily like. The safest notebook I can find is indeed a simple pen and paper because it does not know or care what is being written, it just does it's best regardless of how amazing or horrible the content is.
- SubiculumCode 3y agoIt is interesting to me that these diffusion image models are so much smaller than the LLMs.
- sjm 3y agoThe example images look so bad. Absolutely zero artistic value.
- wongarsu 3y agoFrom a technical perspective they are impressive. The depth of field in the classroom photo and the macro shot. The detail in the chameleon. The perfect writing in very different styles and fonts. The dust kicked up by the donut. The artistic value is something you have to add with a good prompt with artistic vision. These images are probably the AI equivalent of "programmer art". It fulfills its function, but lacks aesthetic considerations. I wouldn't attribute that to the model just yet.
- the_duke 3y agoI'm willing to bet that they are avoiding artistic images on purpose to not get any heat from artists feeling ripped off, which did happen previously.
- robertwt7 3y agoIt’ll be interesting to see what “safety” means in this case given the censorship in diffuser models nowadays. Look what’s happening with Gemini, it’s quite scary really how different companies have different censorship values I’ve had some fair share of frustation with DallE as well when trying to generate weapon images for game assets. Had to tweak a lot of my prompt
- yreg 3y ago> it’s quite scary really how different companies have different censorship values The fact that they have censorship values is scary. But the fact that those are different is better than the alternative.
- declan_roberts 3y agoCan it generate an image of people without injecting insufferable diversity quotas into each image? If so then it’s the most advanced model on the internet right now!
- miohtama 3y agoNo model. Half of the announcement text is “we area really really responsible and safe, believe us.” Kind of a dud for an announcement.
- nextworddev 3y agoThe company itself is about to go run out of money hence the Hail Mary at trying to get acquired
- yreg 3y agoThey raised 110M in October. How much are they burning and how? Training each model allegedly costs hundreds of k.
- haolez 3y agoRewriting the "safety" part, but replacing the AI tool with an imaginary knife called Big Knife: "We believe in safe, responsible knife practices. This means we have taken and continue to take reasonable steps to prevent the misuse of Big Knife by bad actors."
- animex 3y agoUgh, another startup(?) requiring Discord to use their product. :(
- tavavex 3y agoAs far as I know, the Discord thing is only for doing early testing among their community. The full model releases are posted to Hugging Face.
- 13of40 3y ago"we have taken and continue to take reasonable steps to prevent the misuse of Stable Diffusion 3 by bad actors" It's kind of a testament to our times that the person who chooses to look at synthetic porn instead of supporting a real-life human trafficking industry is the bad actor.
- user_7832 3y agoAgree, I think it fundamentally stems from the old conservative view that porn = bad. Morally policing such models is questionable.
- rockooooo 3y agono AI company wants to be the one generating pornographic deepfakes of someone and getting in legal / PR hot water
- seanw444 3y agoWhich is why this should be a much more decentralized effort. Hard to take someone to court when it's not one single person or company doing something.
- mrkramer 3y agoBut what if you flip the things the other way around; deepfake porn is problematic not because porn is per se problematic but because deepfake porn or deepfake revenge porn is made without consent, but what if you give consent to some AI company or porn company to make porn content of you. I see this as evolution of OnlyFans where you could make AI generated deepfake porn of yourself. Another use case would be that retired porn actors could license their porn persona (face/body) to some AI porn company to make new porn. I see big business opportunity in the generative AI porn.
- Cookingboy 3y agoThis is why I think generative AI tech should either be banned or be completely open sourced. Mega tech corporations are plenty of things already, they don't need to be the morality police for our society too.
- GenericPoster 3y agoThe talk of "safety" and harm in every image or language model release is getting quite boring and repetitive. The reasons why it's there is obvious and there are known workarounds yet the majority of conversations seems to be dominated by it. There's very little discussion regarding the actual technology and I'm aware of the irony of mentioning this. Really wish I could filter out these sorts of posts. Hopefuly it dies down soon but I doubt it. At least we don't have to hear garbage about "WHy doEs opEn ai hAve oPEn iN thE namE iF ThEY aReN'T oPEN SoURCe"
- learningerrday 3y agoI hope the safety conversation doesn't die. The societal effects of these technologies are quite large, and we should be okay with creating the space to acknowledge and talk about the good and the bad, and what we're doing to mitigate the negative effects. In any case, even though it's repetitive, there exists someone out there on the Interwebs who will discover that information for the first time today (or whenever the release is), and such disclosures are valuable. My favorite relevant XKCD comic: https://xkcd.com/1053/ https://xkcd.com/1053/
- GenericPoster 3y agoI get that but it just overshadows the technical stuff in nearly every post. And it's just low hanging fruit to have a discussion over. But you're probably right with that comic, I spend so much time reading about ai stuff.
- iterateAutomate 3y agoWhat is with these names haha, Stable Diffusion XL 1.0 and now to Stable Diffusion 3??
- yreg 3y agoThere was 1.0, 1.5, 2.0, XL and now 3.0. Not that weird.
- cchance 3y agoXL was basically an experiment on a 2.1 architecture with some tweaks but at a larger image size... hence the XL but it wasn't really an evolution of the underlying architecture which is why it wasn't 3.0 or even 2.5 it was "bigger" lol
- k__ 3y agoSo, they block all bad actors, but themselves?
- ssalka 3y agoI wonder if this will actually be adopted by the community, unlike SD. 2.0. Many are still developing around SD 1.5 due to its uncensored nature. SDXL has done better than 2.0, but has greater hardware requirements so still can't be used by everyone running 1.5.
- caycep 3y agoare all the model/back ends to Stability products basically available OSS via Ludwig Maximilian University, more or less?
- ummonk 3y agoIt's going to have a restrictive license like Stable Cascade no doubt.
- smrtinsert 3y agoIt's not a restrictive license. They made the model, they trained the base. They're releasing it for consumer use, but not for businesses to effectively re-sell. Makes perfect sense to me.
- ummonk 3y agoThat’s a restrictive license. It’s certainly a reasonable license given the investment they have put into training the model (and Stability’s membership pricing for small companies is, if anything, unreasonably cheap). Nevertheless, it’s frustrating that the industry is fragmenting to a variety of licenses where you have to read the fine print and often licensing information isn’t announced until final release.
- aussieguy1234 3y agoHow does it go with fingers?
- panzi 3y ago503 Service Unavailable welp