3 ms·
To be fair, IIRC OpenAI did release GPT-2-large after some community pressure, and it was somewhat doable for some people to actually train from scratch. GPT-3
by tmabraham 6y ago
To be fair, IIRC OpenAI did release GPT-2-large after some community pressure, and it was somewhat doable for some people to actually train from scratch. GPT-3 is too large, so even if they released it, nobody apart from large companies like Google could do anything with it. If anything, they've made GPT-3 more open than it would have been if they just released the weights.
At least that's my understanding. Feel free to correct me if I'm wrong.
- MiroF 6y agoI'm not really up in arms about OpenAIs decision, but I just don't want people to frame this as "why can't more companies release models like OpenAI does" when in reality the opposite is pretty much true. GPT-2 wasn't really feasible for most actors to train from scratch, otherwise it would have been released by a third party. GPT-3 is technically feasible for a single actor to do inference with, although not really really.
- tmabraham 6y agoI agree with your comment here but I think for the type and size of model, OpenAI has done the best they could with such large models. So I wouldn't say they've "regressed and refused to release its GPT models" they just had to take a different route with GPT models.
- tmabraham 6y agoPeople who disagree and are downvoting, may I ask why?
- Der_Einzige 6y agoThey only released GPT-2 large after other language models came out that were even larger, like T5...
- liuliu 6y agoI don't know. It has 175 billion parameters, thus, about 500GiB for floating point parameters. If these are bfloat16, it is mere 300GiB. We know CPU is about 50 times slower than GPU for transformer models, hence, we are looking at 4 to 8 minutes per inference (parameter loading on demand from SSD takes 200 seconds or so, and probably the bottleneck here). If the parameters can be loaded into memory (seems you have to be on a >= 8-channel machine with unbuffered RAM or with buffered RAM), it will be probably 1 to 2 minutes per inference. Still in the realm of doing inference in homelab territory (barely).
- MiroF 6y ago> We know CPU is about 50 times slower than GPU for transformer models Where's this heuristic from? Seems handy if true.
- liuliu 6y agoIt is from my own experience. But it looks like inference is a bit faster on CPU: https://medium.com/huggingface/benchmarking-transformers-pytorch-and-tensorflow-e2917fb891c2 https://medium.com/huggingface/benchmarking-transformers-pyt...
- hansvm 6y agoThose figures make sense as a rough ballpark -- max FLOPS estimates including impacts from FMA and whatnot roughly correspond to the throughput we would expect for a compute-bound linear algebra problem, and current top of the line consumer grade GPUs are 15-30x better in that metric than equivalent CPUs.
- hansvm 6y agoAside from having to slice and dice things a bit to fit in the accelerator's memory and incurring a higher memory bandwidth cost in doing so, is there any reason a home lab's GPU couldn't be used for GPT-3?
- 6y ago