4 ms·
I think people don't realize how huge these models really are. When they're free, it's pretty cool. But charge an amount where there's actual profit in the pro
by dave_sullivan 4y ago
I think people don't realize how huge these models really are.
When they're free, it's pretty cool. But charge an amount where there's actual profit in the product? Suddenly seems very expensive and not economically viable for a lot of use cases.
We are still in the "you need a supercomputer" phase of these models for now. Something like DALLE mini is much more accessible but the results aren't good enough. Early early days.
- TigeriusKirk 4y agoWhat are the resources at work here? What are the resources needed to train this model? If someone just gave you the model for free, what resources would you need to use it to generate new results?
- binarymax 4y agoIf I had to guess, based on other large models, it’s in the range of hundreds of GBs. It might even be in the TB range. To host that model for fast production SaaS inference requires many GPUs. An A100 has 80GB, so a dozen A100s just to keep it in memory, and more if that doesn’t meet the request demand. Training requires even more GPUs, and I wouldn’t be surprised if they used more than 100 and trained over 3 months.
- judge2020 4y ago> Training requires even more GPUs, and I wouldn’t be surprised if they used more than 100 and trained over 3 months. Based on this blog post where they scale to 7,500 'nodes', they say: > A large machine learning job spans many nodes and runs most efficiently when it has access to all of the hardware resources on each node. So I wouldn't be surprised if they do have a total of 7500+ GPUs to balance workloads between. TO add, OpenAI has a long history of getting unlimited access to Google's clusters of GPUs (nowadays they pay for it, though). When they were training 'OpenAI Five' to play Dota 2 at the highest level, they were using 256 P100 GPUs on GCP[0] and they casually threw 256 GPUs at 'clip' for a short while in January of 2021[1]. As for how they do it, see these posts: https://openai.com/blog/techniques-for-training-large-neural-networks/ https://openai.com/blog/techniques-for-training-large-neural... https://openai.com/blog/triton/ https://openai.com/blog/triton/ 0: https://openai.com/blog/openai-five/ https://openai.com/blog/openai-five/ 1: https://openai.com/blog/clip/ https://openai.com/blog/clip/
- dave_sullivan 4y agoFacebook released over 100 pages of notes a few months ago detailing their training process for a model that is similar in size. Does anyone have a link? I can't seem to find it in my notes, googling links to posts that have been removed or are behind the facebook walled garden. But I seem to remember they were running 1,000+ 32gb GPUs for 3 months to train it and keeping that infrastructure running day-to-day and tweaking parameters as training continued was the bulk of the 100 pages. It is beyond the reach of anybody but a really big company, at least in the area of very large models, and the large models are where all the recent results are. I wish I was more bullish on algorithm improvements meaning you can get better results on less hardware; there will definitely be some algorithm improvements, but I think we might really need more powerful hardware too. Or pooled resources. Something. These models are huge.
- ninjaranter 4y ago> Facebook released over 100 pages of notes a few months ago detailing their training process for a model that is similar in size. Does anyone have a link? Is https://github.com/facebookresearch/metaseq/blob/main/projects/OPT/chronicles/OPT175B_Logbook.pdf https://github.com/facebookresearch/metaseq/blob/main/projec... what you're referring to?
- dave_sullivan 4y agoYes! Thank you! Very good read for anyone interested in the field.
- Ajedi32 4y agoTraining is obviously very expensive, and ideally they'd want to recoup that investment. But I'm curious as to what the marginal cost is to run the model after it's trained. Is it close to 30 images per dollar, like what they're charging now? Or do training costs make up the majority of that price?
- dplavery92 4y agoIn the unCLIP/DALL-E 2 paper[0], they train the encoder/decoder with 650M/250M images respectively. The decoder alone has 3.5B parameters, and the combined priors with the encoder/decoder are the in the neighborhood of ~6B parameters. This is large, but small compared to the name-brand "large language models" (GPT3 et. al.) This means the parameters of the trained model fit in something like 7GB (decoder only, half-precision floats) to 24GB (full model, full-precision). To actually run the model, you will need to store those parameters, as well as the activations for each parameter on each image you are running, in (video) memory. To run the full model on device at inference time (rather than r/w to host between each stage of the model) you would probably want an enterprise cloud/data-center GPU like an NVIDIA A100, especially if running batches of more than one image. The training set size is ~97TB of imagery. I don't think they've shared exactly how long the model trained for, but the original CLIP dataset announcement used some benchmark GPU training tasks that were 16 GPU-days each. If I were to WAG the training time for their commercial DALL-E 2 model, it'd probably be a couple of weeks of training distributed across a couple hundred GPUs. For better insight into what it takes to train (the different stages/components of) a comparable model, you can look through an open-source effort to replicate DALL-E 2.[2] [0] https://cdn.openai.com/papers/dall-e-2.pdf https://cdn.openai.com/papers/dall-e-2.pdf [1] https://openai.com/blog/clip/ https://openai.com/blog/clip/ [2] https://github.com/lucidrains/dalle2-pytorch https://github.com/lucidrains/dalle2-pytorch
- woojoo666 4y ago> This means the parameters of the trained model fit in something like 7GB (decoder only, half-precision floats) to 24GB (full model, full-precision) > you would probably want an enterprise cloud/data-center GPU like an NVIDIA A100, especially if running batches of more than one image. That doesn't seem so bad. looks up price of NVIDIA A100 - $20,000 oh...ok I'll probably just pay for the service then
- fennecfoxen 4y agop4d.24xlarge is only $33/hr! And you get 400 Gbe so it should be quick to load.
- 4y ago
- sinenomine 4y ago> I think people don't realize how huge these models really are. They really aren't that large by the contemporary scaling race standards. DALLE-2 has 3.5B parameters, which should fit on an old GPU like Nvidia RTX2080, especially if you optimize your model for inference [1][2] which is commonly done by ML engineers to minimize costs. With optimized model, your memory footprint is ~1 byte per parameter, and some less than 1 ratio (commonly ~0.2) of all parameters to store intermediate activations. You should be able to run it on Apple M1/M2 with 16GB RAM via CoreML pretty fine, if an order of magnitude slower than on an A100. Training isn't unreasonably costly as well: you can train a model given O(100k)$ which is less than a yearly salary of a mid-tier developer in silicon valley. There is no reason these models shouldn't be trained cooperatively and run locally on our own machines. If someone is interested in cooperating with me on such a project, my email is in the profile. 1. https://arxiv.org/abs/2206.01861 https://arxiv.org/abs/2206.01861 2. https://pytorch.org/blog/introduction-to-quantization-on-pytorch/ https://pytorch.org/blog/introduction-to-quantization-on-pyt...
- gwern 4y agoIt's true that image models are much less of a burden on GPU VRAM than a model like BLOOM where fitting it into a few A100s is ideal, but these diffusion models are a PITA for a ordinary hobbyist in terms of total compute: the CLIP pass over the text input is almost free, but then you feed it into the diffusion model, for one sample you'll be doing 10-100 forward passes (depending on how fancy the diffusion methods are - maybe even 1000 passes if you're using older/simpler ones), and for interactive use, you really want more like 6-9 separate samples; then they have to pass through the upscalers, which are diffusion models themselves and need to do a bunch of forward passes to denoise it. If you do 1 sample in 10s on your 1 consumer GPU, which would be pretty good, 6-9 means a joykilling minute+ wait. And then the user will pick a variation or edit one or tweak the prompt, and start all over again! It's like being back on 16kb dialup waiting for the newest .com to load.
- sinenomine 4y agoAll valid points, of course. As an independent explorer I adapted my workflow to use night's worth of workstation compute to generate a crop of new images from a simple templated prompt language. It also helps a lot to have at least two presets - "exploratory" and "hq", to minimize iteration time and maximize quality of promising prompts. Still, I think optimization of diffusion models for efficient inference isn't yet pushed to the limits. At least if we look at what's available to the public - AFAIK public inference software distributions didn't even quantize their weights.
- spaceman_2020 4y agoHow hard would it be to spin off a variant of this with more focused data models that cater to specific styles or art-types? Like say, a data model only for drawing animals. Or one only for creating new logos?
- mawise 4y agoGenerative networks are worth exploring for randomly creating things in a given category, see this recent HN post about food pictures: https://news.ycombinator.com/item?id=32167704 https://news.ycombinator.com/item?id=32167704