4 ms·
the clipdrop demo doesn't inspire much confidence with how bad the generations are, we're talking 2021 levels, not to mention everything is NSFW somehow, should
by lemoncookiechip 3y ago
the clipdrop demo doesn't inspire much confidence with how bad the generations are, we're talking 2021 levels, not to mention everything is NSFW somehow, should probably work on those filters.
- minimaxir 3y agoI tested a bit and the quality for photorealistic images is surprisingly bad, and definitely worse than LCM and of course normal SDXL. For more artistic images, SDXL Turbo fares better. Unlike normal SDXL, you're required here to use the old-fashioned syntatic sugar like "8k hd" and "hyperrealistic" to align things.
- jauntywundrkind 3y agoAre there any good resources for learning a lot of "syntactic sugar" terms? This is new to me, but I'd love to know more.
- swyx 3y agohttps://github.com/swyxio/ai-notes/blob/main/IMAGE_PROMPTS.md https://github.com/swyxio/ai-notes/blob/main/IMAGE_PROMPTS.m...
- jauntywundrkind 3y ago#LearningInPublic strikes again! I love you swyx!
- swyx 3y agohaha thank you ser. best way to thank me is to start your own LIP practice and pass it on :)
- genewitch 3y agoIt is completely dependent on the model. Civit dot ai has model showcases as well as fine-tune showcases, and you can click any image or press the (i) to see the generation info. Some models like natural language prompts - "draw me a pterodactyl tanning at a beach", some prefer shorthand (danbooru style clip) - "1man, professor, classroom, chalkboard, white_hair, suit", and some work with a mixture of the above as well as the syntactical sugar -"masterpiece, 8k, trending on artstation, space image, a man floating next to a spaceship in space, bokeh, rim lighting, cinematic lighting, Nikon D60, f / 2" Fine-tuning models - LoRA, etc, allow one to convert prompts from one style to another if they wish, but usually it's to compress an idea, style, person, object, etc in to a single "token", so you can work on other aspects of the image. Check out civit AI and you can sort of get an idea of the cargo cultism as well as what sort of keywords actually make a difference.
- dragonwriter 3y ago> Civit dot ai The site you are thinking of is https://civitai.com/ https://civitai.com/ not "civit dot ai".
- dragonwriter 3y agoIn my own testing (using ComfyUI), the best of the "fast gen" techniques for sdxl is using the Turbo model [0], but using the LCM sampler with the sgm_uniform scheduler (which is normal for LCM) with it, and running it up to 4-10 iterations instead of just one. I think StabilityAI demos are using Euler A with the normal base scheduler, and running a single iteration (which is cool for a max-speed demo, and its awesome for that speed, but its leaving a lot of quality on the table that you can get with a few more iterations especially with the LCM/sgm_uniform sampler/scheduler combo.) Bumping CFG up slightly helps, too (but I think adds another performance hit, because I think the demos are running at CFG 1, which AIUI disables CFG and reduces computations per iteration.) > Unlike normal SDXL, you're required here to use the old-fashioned syntatic sugar like "8k hd" and "hyperrealistic" to align things. That's not "syntactic sugar", and its not particular my experience that it is needed with sdxl turbo. [0] actually, differencing the base sdxl model from the turbo model to get a "turbo modifier", and then combining that with a good SDXL-based checkpoint, because StabilityAI's base models are pretty ho-hum compared to decent community checkpoints derived from them, but that is kind of a peripheral issue.
- Jackson__ 3y agoYeah, it looks like the enshittification of StabilityAI is in full force by now. Especially considering the continually worse licensing. I expect if they ever manage to release an image gen model that's an objective improvement, lets say 80% as good as dalle3, it will be subscription API only.
- BoorishBears 3y agoAre you serious? I'm using Stability in production: they kept their SDXL beta model which was capable of SDXL 1.0 level prompt adherence at a fraction of the cost up for months after was reasonable for a one-off undocumented beta, and it was a huge boon to my product. Then a few weeks back they went and quietly cut costs to 1/5th or so what they were for SDXL and released a model that produced similar quality outputs to SDXL for my specific usecase in a fraction of the time (SD 1.6) They're on fire as far as I'm concerned, just quietly making their product cheaper and faster. — Also Dalle 3 is in a very awkward place for programmatic access, so awkward I wouldn't call them competitive to SD for many usecases: It's got a layer of prompt interference baked in, it's expensive, latency is not very consistent. Text is a cool trick but it's still not reliable enough to expose as a core part of the generation for an end user.
- ShamelessC 3y agoSounds like they’re doing the same thing OpenAI is doing. Claiming to favor open models but the reality is they’re pumping growth by reducing costs and this lowering prices. They want a massive chunk of this new market, all of it if they can get it. Their perceived valuation then becomes a matter of how many eyeballs they have looking at segments of their website to advertise to, or how many data points they can collect on their users to sell to advertisers. It’s unlikely they can capture the whole market and still make a chunky enough profit to satisfy investors if they also intend to keep prices high enough without needing to resort to enshitification.
- BoorishBears 3y agoThis would be a lot more pithy if it weren't in the comment section of a post that showcases exactly how they were likely able to make 1.6 cheaper, and open sources the underlying tech. There couldn't be a more perfect rebuttal to this theory than the post you decided to leave it under.