6 ms·
Been using it for a couple of hours and it seems it’s much better at following the prompt. Right away it seems the quality is worse compared to some SDXL models
by obviyus 3y ago
Been using it for a couple of hours and it seems it’s much better at following the prompt. Right away it seems the quality is worse compared to some SDXL models but I’ll reserve judgement until a couple more days of testing.
It’s fast too! I would reckon about 2-3x faster than non-turbo SDXL.
- sorenjan 3y agoHow much VRAM does it need? They mention that the largest model uses 1.4 billion parameters more than SDXL, which in turn need a lot of VRAM.
- adventured 3y agoThere was a leak from Japan yesterday, prior to this release, and in that it was suggested 20gb for the largest model. This text was part of the Stability Japan leak (the 20gb VRAM reference was dropped in the release today): "Stages C and B will be released in two different models. Stage C uses parameters of 1B and 3.6B, and Stage B uses parameters of 700M and 1.5B. However, if you want to minimize your hardware needs, you can also use the 1B parameter version. In Stage B, both give great results, but 1.5 billion is better at reconstructing finer details. Thanks to Stable Cascade's modular approach, the expected amount of VRAM required for inference can be kept at around 20GB, but can be reduced even further by using smaller variations (as mentioned earlier, this (which may reduce the final output quality)."
- sorenjan 3y agoThanks. I guess this means that fewer people will be able to use it on their own computer, but the improved efficiency makes it cheaper to run on servers with enough VRAM. Maybe running stage C first, unloading it from VRAM, and then do B and A would make it fit in 12 or even 8 GB, but I wonder if the memory transfers would negate any time saving. Might still be worth it if it produces better images though.
- adventured 3y agoIf it worked I imagine large batching could make it worth the load/unload time cost.
- weebull 3y agoShouldn't be a reason you couldn't do a ton of Layer C work on different images, and then swap in Layer B.
- Filligree 3y agoSequential model offloading isn’t too bad. It adds about a second or less to inference, assuming it still fits in main memory.
- sorenjan 3y agoSometimes I forget how fast modern computers are. PCIe v4 x16 has a transfer speed of 31.5 GB/s, so theoretically it should take less than 100 ms to transfer stage B and A. Maybe it's not so bad after all, it will be interesting to see what happens.
- whywhywhywhy 3y agoIf you're serious about doing image gen locally you should be running a 24GB card anyway because honestly Nvidia's current generation 24GB is the sweet spot price to performance. 3080 ram is laughably the same as the 6 year old 1080Ti and 4080 ram is only slightly more at 16 and costs about 1.5 times the 3090 second hand. Any speed benefits of the 4080 are gonna be worthless the second it has to cycle a model in and out of ram anyway vs the 3090 in image gen.
- weebull 3y ago> because honestly Nvidia's current generation 24GB is the sweet spot price to performance How is the halo product of a range the "sweet spot"? I think nVidia are extremely exposed on this front. The RX 7900XTX is also 24GB and under half the price (In UK at least - £800 vs £1,700 for the 4090). It's difficult to get a performance comparison on compute tasks, but I think it's around 70-80% of the 4090 given what I can find. Even a 3090, if you can find one, is £1,500. The software isn't as stable on AMD hardware, but it does work. I'm running a RX7600 - 8GB myself, and happily doing SDXL. The main problem is that exhausting VRAM causes instability. Exceed it by a lot, and everything is handled fine, but if it's marginal... problems ensue. The AMD engineers are actively making the experience better, and it may not be long before it's a practical alternative. If/When that happens nVidia will need to slash their prices to sell anything in this sphere, which I can't really see themselves doing.
- liuliu 3y agoShould use no more than 6GiB for FP16 models at each stage. The current implementation is not RAM optimized.
- sorenjan 3y agoThe large C model uses 3.6 billion parameters which is 6.7 GiB if each parameter is 16 bits.
- liuliu 3y agoThe large C model have fair bit of parameters tied to text-conditioning, not to the main denoising process. Similar to how we split the network for SDXL Base, I am pretty confident we can split non-trivial amount of parameters to text-conditioning hence during denoising process, loading less than 3.6B parameters.
- brucethemoose2 3y agoWhat's more, they can presumably be swapped in and out like the SDXL base + refiner, right?
- kimoz 3y agoCan one run it on CPU?
- ghurtado 3y agoYou can run any ML model on CPU. The question is the performance
- rwmj 3y agoStable Diffusion on a 16 core AMD CPU takes for me about 2-3 hours to generate an image, just to give you a rough idea of the performance. (On the same AMD's iGPU it takes 2 minutes or so).
- OJFord 3y agoEven older GPUs are worth using then I take it? For example I pulled a (2GB I think, 4 tops) 6870 out of my desktop because it's a beast (in physical size, and power consumption) and I wasn't using it for gaming or anything, figured I'd be fine just with the Intel integrated graphics. But if I wanted to play around with some models locally, it'd be worth putting it back & figuring out how to use it as a secondary card?
- rwmj 3y agoOne counterintuitive advantage of the integrated GPU is it has access to system RAM (instead of using a dedicated and fixed amount of VRAM). That means I'm able to give the iGPU 16 GB of RAM. For me SD takes 8-9 GB of RAM when running. The system RAM is slower than VRAM which is the trade-off here.
- OJFord 3y agoYeah I did wonder about that as I typed, which is why I mentioned the low amount (by modern standards anyway) on the card. OK, thanks!
- mat0 3y agoNo, I don't think so. I think you would need more VRAM to start with.
- vergessenmir 3y agoI'll take prompt adherence over quality any day. The machinery otherwise isn't worth it i.e the controlnets, openpose, depthmaps just to force a particular look or to achieve depth. Th solution becomes bespoke for each generation. Had a test of it and my option is it's an improvement when it comes to following prompts and I do find the images more visually appealing.
- stavros 3y agoCan we use its output as input to SDXL? Presumably it would just fill in the details, and not create whole new images.
- RIMR 3y agoI was thinking that exactly. You could use the same trick as the hires-fix for an adherence-fix.
- emadm 3y agoYeah chain it in comfy to a turbo model for detail
- Filligree 3y agoA turbo model isn't the first thing I'd think of when it comes to finalizing a picture. Have you found one that produces high-quality output?
- dragonwriter 3y agoFor detail, it'd probably be better to use a full model with a small number of steps (something like KSampler Advanced node with 40 total steps, but starting at step 32-ish.) Might even try using the SDXL refiner model for that. Turbo models are decent at low-iteration-decent-results, but not so much at adding fine details to an mostly-done image.