6 ms·
How much would it cost to train something like this? Is there even a good dataset for it?
by bufferoverflow 4y ago
How much would it cost to train something like this? Is there even a good dataset for it?
- Yenrabbit 4y agoThere is a dataset of 5 billion image-text pairs (laion-5b) scraped by various parties. This can then be filtered and used to train these models. Cost is a bit of an issue but there are orgs that have provided compute for open model training. And Imagen is nice because the text encoder part is already available and doesn't need more training, so it would just be the diffusion model components being trained. I'd guess we'll see a biggish training run starting in a few weeks.
- visarga 4y agoI hope so. It's a bit cruel to show off and then lock it away.
- dharma1 4y agoWhen you say a bit of an issue, how much are we talking?
- Yenrabbit 4y agoFour or five figures I'd guess? I'm not clued up on costs/performance for TPU stuff to give a better estimate, but guessing at a week on a 256 TPU pod, call it $30k?
- sailingparrot 4y agoYou are off by an order of magnitude at least. 256 TPUs-v4 (not pod), would cost you around 20k$/day. They actually used 512 TPUs (256 for base model + 128 for each of the two superresolution models). Assuming an average training time of 1 week as you said, that gives us about 280k$. It's also most likely trained for longer than a week, the base model for Dalle-2 was trained for 100-200k GPU hours, so between 2-4x longer than that, we can guess this is roughly similar. You also never successfully train everything first try, so all in all, to replicate this work just from the paper, we are talking about at least 500k$.
- rockemsockem 4y agoWhile that is what they did, they also used a batch size of 2048 while training. This is just to speed training up, not a hard requirement. It's easy for Google to justify more money on compute to save engineer iteration loops. I'll have to read the paper for more details, but it would almost certainly cost less (and take longer) to train a model like this in a more resource constrained situation than Google faces .
- sailingparrot 4y agoIncreasing batch size does not increase the cost of your training. The opposite actually: With bigger batch size (to an extent), models tend to converge slightly faster so you need less GPU hours. As for the rest, training for 1h with batch size of 2048 on 10 TPUs, or training for 8h with batch size of 256 on 1 TPU has the exact same cost, the cost is just spread over a longer time.
- tomatowurst 4y agoCouldn't a bunch of us shell out $5000~$50,000 and do this ourselves? Create a non-profit shell corporation outside US jurisdiction, issue shares, raise funds and open source the result? The shares would simply be votes towards future training dataset endeavors as no profit would be booked here. Say you buy 5000 out of 500,000 shares, that would give you 1% voting power in what dataset to train.
- joshvm 4y agoTF Research Cloud. Especially if multiple people could use it with eg model checkpoints: > 5 on-demand Cloud TPU v3 devices, 5 on-demand Cloud TPU v2 devices, and 100 preemptible Cloud TPU v2 devices for free for 30 days So up to 7k hours on demand and 70k pre-emptible
- bergenty 4y agoWho is going to be liable for the flood of child pornography that comes out of that setup?
- 4y ago
- srcreigh 4y agoWho will do the training? Will the training results be made available publicly? why wont it start until a few weeks from now?
- webmaven 4y agoAre there prefiltered derivatives of Laion-5B available? I can imagine various contraindicated categories you might want to avoid entirely, as well as biases you might want to adjust for by balancing classes in the data (5 billion images gives you a lot of room to balance the dataset).
- dirtyid 4y agoI feel like the fact big porn hasn't poached talent and jumped all over this suggest at least 10s of millions. That said some for profit no-rules deepfake service for disinformation and illegal content has to be in the works.
- tomatowurst 4y agoThere's a company in Montreal that makes that in a month and also has access to copious amount of said datasets on their servers. It may or may not be that they already have engineers on it. We have no way of knowing since its a private company