4 ms·
This is a nice release, but the title is a bit misleading as the released sizes (1.3B and 2.7B parameters) do not yet compare to the size of GPT-3 (175B), but r
by ve55 6y ago
This is a nice release, but the title is a bit misleading as the released sizes (1.3B and 2.7B parameters) do not yet compare to the size of GPT-3 (175B), but rather GPT-2 (1.5B) instead (although future releases may have significantly more!).
Edit: title improved, thank you!
- nl 6y agoYeah. They say they are doing a 10B release soon[1]. I suspect they have run into training issues since they are moving to a new repo[2] [1] https://twitter.com/arankomatsuzaki/status/1373732646811967490 https://twitter.com/arankomatsuzaki/status/13737326468119674... [2] https://github.com/EleutherAI/gpt-neox/ https://github.com/EleutherAI/gpt-neox/
- chillee 6y agoIt's more about hardware - these models were trained on TPUs, while GPT-NeoX is being trained on GPUs graciously provided by Coreweave.
- orra 6y agoAny idea what the required GPU time would cost (if not donated)? Is GPT-3 just a commodity soon?
- Voloskaya 6y ago~4M$ per full training give or take.
- teruakohatu 6y agoThe number thrown around for gpt-3 is $4.6 million, but I am not sure where that figure originates.
- minimaxir 6y agoIt was a number tossed around by a GPU hosting provider, based on their own costs: https://lambdalabs.com/blog/demystifying-gpt-3/ https://lambdalabs.com/blog/demystifying-gpt-3/ The reality is that GPT-3 was likely "free" to train on Azure, as Microsoft has provided a lot of resources to OpenAI.
- exikyut 6y agoIf this is true, I wonder what sort of social capital transactional exchange is going on instead.
- minimaxir 6y agoWith training improvements such as DeepSpeed, the GPU costs will likely be substantially lower than what was available at the time OpenAI trained GPT-3. Still not free, though. The hard part with GPT-3 is it's big enough to make it difficult to actually deploy.
- stellaathena 6y agoOur current estimate is that it requires between 2000 and 4000 V100 months.
- pizza 6y agoFixed title to reflect that, thanks
- ve55 6y agoI would perhaps change 'GPT-3' to just say 'GPT' instead, as a more salient fix.
- nl 6y agoThis is incorrect. It's the GPT-3 model architecture and optimisations, and uses training techniques similar to GPT-3.
- ve55 6y agoThank you, I've rephrased a few things to improve the wording with respect to this.
- stellaathena 6y agoGPT-3 isn't a single model. It's a model architecture that is very closely followed by GPT-Neo. The 2.7B model is the exact same size as something OpenAI sells under the label "GPT-3"
- Dylan16807 6y agoIs GPT-2's architecture any different?
- stellaathena 6y agoNot hugely, but yes. I tend to think of GPT as a style of architecture with consistent themes and major features, but varying minor features and implementation details. Off the top of my head, I believe the most important difference is that GPT-3 alternates global and local attention while GPT-2 is all global attention. The two published GPT-Neo models follow GPT-3's lead but the repo lets the user pick whether to use global or local attention layers.