6 ms·
Zero-3 Offload: Scale DL models to trillion parameters without code changes
- bionhoward 6y agoplease hook this up to Jax!
- joshlk 6y agoGPT-NeoX is an example project that is using deepspeed and Zero-3 offloading. The wider project intend to train a GPT-3 sized model and release it freely to the world. https://github.com/EleutherAI/gpt-neox https://github.com/EleutherAI/gpt-neox
- ma2rten 6y agoIt seems like Zero-3 doesn't work for them: https://github.com/EleutherAI/gpt-neox/issues/171 https://github.com/EleutherAI/gpt-neox/issues/171
- dqpb 6y agoDid you even read through the issue? I don't see anything that indicates it won't work.
- ma2rten 6y agoYes, I did. The last comment is a traceback and an explanation what would have to be done to fix it.
- joshlk 6y agoLooks like they got it working recently https://github.com/EleutherAI/gpt-neox/pull/178 https://github.com/EleutherAI/gpt-neox/pull/178
- stellaathena 6y agoHi! I’m the one who wrote this code. My ZeRO-3 implementation is currently not working, but I’ve spoken with DeepSpeed devs and they’ve explained to me what I’ve been doing wrong. I haven’t had time to implement the fix but I don’t see any reason to assume it won’t work. https://github.com/microsoft/DeepSpeed/issues/846 https://github.com/microsoft/DeepSpeed/issues/846 Also, the specific problem described in that Issue was due to a bug I found in DeepSpeed that has since been corrected.
- FL33TW00D 6y agoHuggingface has been working on implementing this into their library, and it has some pretty amazing effects on the size of models you can train on a simple Colab. https://huggingface.co/blog/zero-deepspeed-fairscale https://huggingface.co/blog/zero-deepspeed-fairscale
- stephenroller 6y agoSupport for this was also added to [Fairscale](https://fairscale.readthedocs.io/en/latest/ https://fairscale.readthedocs.io/en/latest/) and [Fairseq](https://github.com/pytorch/fairseq https://github.com/pytorch/fairseq) last week. In particular, the Fairscale implementation can be used in any pyotrch project without requiring the use of the Deepspeed trainer.
- diptanu 6y agoWhat are the relevant commits in Fairseq for this? I couldn't figure out the changes by looking at the commits from last week.
- stephenroller 6y agohttps://github.com/pytorch/fairseq/pull/3331 https://github.com/pytorch/fairseq/pull/3331 and https://github.com/pytorch/fairseq/pull/3327 https://github.com/pytorch/fairseq/pull/3327
- bevenky 6y agoThis is also being added to pytorch https://github.com/pytorch/pytorch/pull/46750 https://github.com/pytorch/pytorch/pull/46750
- minimaxir 6y agoI don't think that's the Stage 3 announced in this blog post, but it's def a framework for it.
- andrewprock 6y agoHow much data do you need to mitigate the risk of over fitting a trillion parameter model?
- gwern 6y agoYou ideally need ~500GB of text, or so. EleutherAI's The Pile was designed to be just big enough to fit a 1t GPT efficiently, and you can get the various scaling curves out of the OA-related scaling papers. (You want the amount of data that fits into a single epoch, because if you reuse data, you get less bang for the FLOPs buck, and FLOPS constraints are right now much more binding than data or model size.)
- andrewprock 6y agoThis feels off by a couple of orders of magnitude, unless a significant number of the parameters are not independent.
- gwern 6y agoIt's quite amusing. The standard statistical theory does not work at all in estimating data vs model size, and the bounds are all vacuously large. It's a very active area of research, understanding why models act so simple when overparameterized and coming up with real measures of model complexity. Lots to read there if you are interested in such things.
- andrewprock 6y agoThat just means that the parameters are not independent.
- gwern 6y agoBut you can fit randomly-generated labels!
- 6y ago
- dataangel 6y agoELI5? All this techno babble just sounds like "it's faster because we optimized it". What are the nontrivial, new fundamental tricks?
- jiofih 6y agoThird paragraph or so in the overview: > ZeRO removes the memory redundancies across data-parallel processes by partitioning the three model states (optimizer states, gradients, and parameters) across data-parallel processes instead of replicating them. By doing this, it boosts memory efficiency compared to classic data-parallelism while retaining its computational granularity and communication efficiency
- dataangel 6y agoYeah that would be the techno-babble. I've been working on a machine learning pipeline for 6 years and I still have no idea what this means.
- jiofih 6y agoYou can read the paper here: https://arxiv.org/abs/1910.02054 https://arxiv.org/abs/1910.02054
- cambalache 6y agoThe product is obviously not for you but for clueless PHBs who want the "latest and best" for the team so those useless ML engineers can finally put his brilliant idea in production with a less than 1% prediction error.
- zachthewf 6y agoIt doesn't sound like techno-babble to me. They've distributed storage across nodes rather than replicating on each node, hence the model size is now scalable with number of nodes rather than being limited to what could be stored on a single node.
- 6y ago
- ansk 6y agoQuestion for someone knowledgable about this: if I have a model which is large -- but small enough that I can fit a single training example on GPU -- does this approach offer speedups compared to simple gradient accumulation? Or is this only useful for models which are so large that the model parameters themselves are overwhelming GPU memory?
- vladf 6y agoAlternatively, one could get rid of the memory used by optimizers entirely by switching to vanilla SGD. I haven’t tried this on transformers and maybe that’s what breaks down here but in “classic” supervised settings I’ve found SGD with schedule tuning just as fast as Adam.
- gwern 6y agoSGD doesn't work on large Transformers, no. You need something like AdamW.
- The_rationalist 6y agoMish is generally superior to RadamW https://lessw.medium.com/meet-mish-new-state-of-the-art-ai-activation-function-the-successor-to-relu-846a6d93471f https://lessw.medium.com/meet-mish-new-state-of-the-art-ai-a...
- dron57 6y agoThat's cool, but Mish is an activation function while SGD and AdamW are optimizers. Apples and oranges.
- mchusma 6y agoThis is super impressive. I could not figure out for a while who exactly was running this project, but it looks like its Microsoft. Great work!
- The_rationalist 6y agoSee also zeroth order backpropagation which allows 300X faster training while not reducing throughput that much https://arxiv.org/abs/2011.08895 https://arxiv.org/abs/2011.08895 How much zero-3 affect accuracy? See also https://github.com/microsoft/fastformers https://github.com/microsoft/fastformers
- singhrac 6y agoFor those searching, DeepSpeed is implemented as a set of C++/CUDA extensions on top of PyTorch (compiled using their JIT).
- alphagrep12345 6y agoSimple 10 min overview/tutorial (official) if someone is interested - https://www.youtube.com/watch?v=ovQC7FqXHXk https://www.youtube.com/watch?v=ovQC7FqXHXk