6 ms·
It’s really crazy how Stable Diffusion seems to be very on par with DALL-E and you can run it on “most” hardware. Is there an equivalent for GPT-3? I don’t even
by syntaxing 4y ago
It’s really crazy how Stable Diffusion seems to be very on par with DALL-E and you can run it on “most” hardware. Is there an equivalent for GPT-3? I don’t even think I can run the 2M lite GPT-J on my computer…
- xor99 4y agoThis is the killer aspect of it. Running an image in a <5 mins on a Mac is amazing when you consider the alternatives atm.
- acapybara 4y agoFull GPT-6B can run if you have 22gb ram (CPU or GPU depending on where you run it). Also can run an 8 bit quantized version pretty easily. This takes ~6gb RAM. The results seem far off from GPT-3 but apparently it can get good results when fine tuned. Bigger models like OPT 66B can run on cloud machines (or a really big local system) OPT 175B weights are not open but can be applied for. 175B would require something like 500GB RAM if not quantized. That's a lot, but it's possible to build that locally if you have a couple 10's of thousands of dollars. Wait a few years and 175B on a GPU will be no problem.
- boppo1 4y agoWhat does 'quantized' mean in this context?
- acapybara 4y agoBasically stuff a 32 bit value into an 8 bit value (and lose precision). Apparently it doesn't affect the results significantly. More info: https://github.com/huggingface/transformers/pull/17901 https://github.com/huggingface/transformers/pull/17901
- ManuelKiessling 4y agoTangential: I've set up a Discord Bot that turns your text prompt into images using Stable Diffusion. You can invite the bot to your server via https://discord.com/api/oauth2/authorize?client_id=1013373043625705513&permissions=117760&scope=bot%20applications.commands https://discord.com/api/oauth2/authorize?client_id=101337304... Talk to it using the /draw Slash Command. It's very much a quick weekend hack, so no guarantees whatsoever. Not sure how long I can afford the AWS g4dn instance, so get it while it's hot. Oh and get your prompt ideas from https://lexica.art https://lexica.art if you want good results. PS: Anyone knows where to host reliable NVIDIA-equipped VMs at a reasonable price?
- olladecarne 4y agoOne thing I noticed is that on GCP if you create a a2-ultragpu (Nvidia a100 80gb) and you select a spot instance, the price estimate goes down to $0.33 hourly ($240/m) which sounds really good if it's not a mistake. I was wondering if you could then turn a single A100 into 7 GPUs using Multi-instance GPUs. So on an 80gb one you get 7 10GB GPUs (can't have 8 due to yield issues on those cards). I'm pretty sure that will run much slower than on the full instance, but not 7x slower so if you're running a larger service at scale this could be an option to parallelize things. If someone is able to get that running please let me know how it performs. The next thing I considered was just buying up a ton of 3060 12gb cards (saw a few new ones for $330) and just hosting a server from my house. This might be a good option if you don't care about speed but care about throughput. RTX 3090s are also decent in terms of price per iteration of Stable Diffusion. If you want to build a fast service like Dreamstudio I think it's the only option to be able to do it at a reasonable price. If you want to host these in the cloud using consumer RTX cards, you'll have to go with less reputable hosts since Nvidia doesn't allow it. I don't want to name any since I can't vouch for them, but there are some if you search. The cheapest option will be to buy them and host it yourself. I'm still researching what the best price/performance is for hosting this so if you have any findings please share.
- rexreed 4y agoI'm experimenting with your Discord bot right now. It would be great to have a command that shows where your processes are currently in the queue or maybe the discord bot can update on queue position.
- ManuelKiessling 4y agoGood idea, I'll look into it.
- rexreed 4y agoI submitted 2 /draw requests with prompts, got quoted a time 15-30 min for first one and then 17-34 for 2nd, submitted about 5 minutes apart but it's been now past the upper limit of the quoted time without any results. I'm assuming that the image generation has failed or perhaps the bot got stuck. Having some way of knowing would be helpful.
- juliensalinas 4y agoI worked on the Stable Diffusion and GPT-J integrations on NLP Cloud (https://nlpcloud.com/ https://nlpcloud.com/). Both can be used in FP16 without any noticeable quality drop (in my opinion). Stable diffusion requires 7GB of VRAM on a Tesla T4 GPU. GPT-J requires 12GB of VRAM (but if you really try to use the 2048 tokens context, the VRAM will go up and reach something like 20GB of VRAM).
- pdntspa 4y agoPersonally I don't find it on-par or even close to DALL-E... stylistically its output is a lot more plain (Midjourney does really well here) and it can't handle complicated prompts well (it will pick the one thing in the prompt it does know about and run with it, ignoring all else) Plus, there are huge gaps in training. Ask it to draw something simple, like "a penis" and you get nightmare fuel....
- deleted 4y ago[deleted]
- boppo1 4y agoDoes DALL-E let you output penises? I thought openAI was forbidding many 'unseemly' prompts.
- pjgalbraith 4y agoIt definitely doesn't let you and you may get permanently banned for trying.
- macrolime 4y agoGPT-3 isn't really all that optimized in terms of size. Later studies have shown that you don't need that many parameters to get the same results, so it should be possible to train a model that could run at least on something like an RTX 4090 Ti with 48GB ram.
- karolist 4y agoI have 64GB RAM and nVidia with 24GB vmem, which projects could be the limit I can run locally?
- planetsprite 4y agoStable Diffusion seems hyper-trained on digital art and faces. Dall-e feels a lot more "intelligent" and can create a far greater and more comprehensive diversity of images from different prompts.