3 ms·
There are already open source LLMs with comparable parameter counts (Facebook's OPT-175B, BLOOM), but you'll need ~10x A100 GPUs to run them (which would cost ~
by antimatter15 4y ago
There are already open source LLMs with comparable parameter counts (Facebook's OPT-175B, BLOOM), but you'll need ~10x A100 GPUs to run them (which would cost ~$100K+).
I suspect a big part of why stable diffusion managed to consume so much mindshare is that it can run on ordinary consumer hardware. On that point, I would be excited about an open-source RETRO (https://arxiv.org/pdf/2112.04426.pdf https://arxiv.org/pdf/2112.04426.pdf) model with comparable performance to GPT-3 that could run on consumer hardware with an NVMe SSD.
- rafaelero 4y agoThey aren't as good as davinci-003, though. There is no open source model competitive with GPT-3.5 yet.
- lagrange77 4y agoGPT-J-6B is said to be also of comparable performance in certain areas to GPT-3 and can be run (inference) on a RTX3090. https://huggingface.co/EleutherAI/gpt-j-6B https://huggingface.co/EleutherAI/gpt-j-6B
- Roark66 4y agoOne can in theory run even 175B bloom with just a modern multicore CPU, 32gb of RAM and 2TB of nvme storage. There is a library called accelerate witch with some slight modifications allows one to run in cpu/storage only mode models that don't fit in memory. Of course it takes a looong time to do inference, but one can at least have a taste. The biggest bloom I personally have run on cpu only in this fashion is 7B. It requires 4x7B of RAM plus some. On my hardware it tends to use all 32GB RAM and about ~4GB of storage during inference. At the moment I believe there is still a limitation of the smallest layer fitting in memory at once. This is why I haven't tried bigger bloom, but I believe there are ways to overcome it. Once this problem is resolved one should be able to use the same tech to use GPUs with less vram (like my 2070 with 8GB) for parts of larger models.