3 ms·
The closest open source contender is BLOOM: https://huggingface.co/bigscience/bloom https://huggingface.co/bigscience/bloom. It has an almost identical architec
by spi 4y ago
The closest open source contender is BLOOM: https://huggingface.co/bigscience/bloom https://huggingface.co/bigscience/bloom. It has an almost identical architecture to GPT-3 (hence, to ChatGPT), and in particular the same number of parameters (175B). It was also trained on a similar amount of data, as far as we can know. Still, it's not like you can just "download it and run it", even just to _load_ the model into memory you need ~400GB of memory, to run it at any decent speed you need a lot of GPUs, so it's not really like consumer hardware. And the process to train it cost about 2 to 4 million $, so replicating it is definitely not for everybody. But also not just for "big corporations"...
- throwaway743 4y agoSaw this on here the other day. Uses BLOOM. Not as good as chatGPT, but it's something http://chat.petals.ml/ http://chat.petals.ml/
- espadrine 4y agoIn terms of quality, I think BLOOMZ, or mT0, are the best open-source ones. The non-finetuned BLOOM does not appear favorably (in English) compared to GLM or OPT, which both have published weights: https://crfm.stanford.edu/helm/v0.1.0/?group=mmlu https://crfm.stanford.edu/helm/v0.1.0/?group=mmlu and Flan-T5 is above OPT-IML: https://arxiv.org/pdf/2212.12017.pdf https://arxiv.org/pdf/2212.12017.pdf > Is the future going to be controlled by big corporations who own the models themselves? On this subject, there is an effort stemming from BigScience to build an open, distributed inference network, so that people that don’t have enough GPUs at home can contribute theirs and get text generation at one word per second: https://github.com/bigscience-workshop/petals#how-does-it-work https://github.com/bigscience-workshop/petals#how-does-it-wo...
- metadat 4y agoGetting a server with > 400GB of RAM and a heap of GPUs can be done for less than $6,000 - $10,000 if you're scrappy. Not cheap, but also not out of reach for individuals.
- spi 4y agoI don't think that figure is correct, you need a "good heap" of GPUs, not just anything... in particular, even just to run inference, you need at least 400 GB of GPU memory, not just RAM. You can't just plug a dozen "cheap" GPUs and call it a day, because if I remember correctly consumer GPUs have at most 32GB of RAM each. Hence you'd need at least 12 of those top-tier GPUs (which certainly don't come at $500 a piece). Probably more, because you can't trivially split weights across GPUs so perfectly (you probably have to put an integer number of layers on each GPU). In practice these models are typically run using top-tier A100 GPUs, which apparently is the cheapest thing you can do at scale: https://forum.effectivealtruism.org/posts/foptmf8C25TzJuit6/gpt-3-like-models-are-now-much-easier-to-access-and-deploy https://forum.effectivealtruism.org/posts/foptmf8C25TzJuit6/.... It looks like you can get away with just $10/hour, but I'm not sure I believe it. In one hour you can roughly generate 6 million English words this way, that's quite cheap. But if you want to own the full hardware, then it's quite more expensive. You need 8 of those A100 GPUs, which come at $32k a piece, so you're in the ballpark of > $300k to build the server you need. Then there's of course running costs, these GPUs burn 250W a piece, plus the rest of the server we're at about 3kW power. That's not much, maybe $0.50/hr, plus maybe another $1/hr to cool the room it's in, depending on where it is (and the season, I guess in winter a fan might suffice, it's about as powerful as a couple small electric heaters). So with an upfront expense of > $300k, you're maybe down from $10/hr to $1.5/hr, saving something like $8.5/hr, which is $6k / month (minus the rent of whatever place you put the server in). All in all, it's definitely feasible for a small start up as well, but not very much for an individual.
- metadat 4y agoGot it, thanks for the information! I hadn't known it was all VRAM for model serving.