4 ms·
The reason is that Dall-E 2 type models are small and can run on a wide class of commodity hardware. This makes them very accessible which means a large number
by Vetch 4y ago
The reason is that Dall-E 2 type models are small and can run on a wide class of commodity hardware. This makes them very accessible which means a large number of people can contribute.
Large language models gain key capabilities as they increase in size: more reliable fact retrieval, multistep reasoning and synthesis, complex instruction following. The best publicly accessible is GPT-3 and at that scale you're looking at hundreds of gigabytes.
Models able to run on most people's machines fall flat when you try to do anything too complex with them. You can read any LLM paper and see how the models increase in performance with size.
The capabilities of available small models have increased by a lot recently as we've learned how to train LLMs but a larger model is always going to be a lot better, at least when it comes to transformers.
- dash2 4y agoIs there no way to do kind of split-apply-combine with these models? So you could train GPT@home?
- Vetch 4y agoFor inference, the best models are so large they won't fit in System RAM. GPT@home is not going to make a difference in that scenario. For training such large models, data parallelism is no longer sufficient and tensor/pipeline parallelism is required. The problem is communication bottlenecks, differing device/network speeds and massive data transfer requirements become serious enough issues to kill any naive distributed training across the internet approach. Deep learning companies use fancy 100Gbps+ connections, do kernel hacking and use homogeneous hardware and it's still a serious challenge. There is no incentive for them to invest in something like GPT@home. But it's not impossible and there's some research being done in the area. Although, it'll be a while until a GPT@home approach becomes a ready alternative. See https://arxiv.org/abs/2206.01288 https://arxiv.org/abs/2206.01288 and their recent GPT-JT test for more. Another development would be for networks to become more modular.
- nullc 4y agoRam isn't terribly expensive, it's not unreasonable to have 1 or 2 TB of ram. 1TB costs about $3500 as 64GB dimms. (some of my 4u hosts have 96 ddr4 sockets too... though 6tb of ram is getting a little pricey. :)) > use fancy 100Gbps+ connections, you can pick up 100gbps mellanox nics on ebay for $50 on a good day, $200 whenever. If you're only connecting up two or three hosts you can just use multiport cards and a couple dac cables, rather than a switch. I suspect for inference though there is a substantial locality gain if you're able to batch a lot of users into a single operation, since you can stream the weights through while applying them to a bunch of queries at once. But that isn't necessarily lost on a single user, it would be nice to see a dozen distinct completions at once.
- doctoboggan 4y agoIf you're opinion, what is the best model I can run on my M1 MBP with 64gb memory and 32 GPU cores?
- Vetch 4y agoFor practical tasks, I would like to say FlanT5 11B which is 45GB but my experience is if you're using huggingface the usual way, it can initially take up to 2x the memory of the model to load. GPT-JT was released recently and seems interesting but I haven't tried it. If you're focused on scientific domain and want to do Open book Q/A, summarization, keyword extraction etc. Galactica 6B parameter version might be worth checking out. If our main language is not English one of the mt0 models might be worth a try https://huggingface.co/bigscience/mt0-xl https://huggingface.co/bigscience/mt0-xl These models are distinguished by being able to follow relatively complex natural language instructions and examples without needing to be finetuned.
- zaptrem 4y agoI'm able to run a 22b parameter GPT-Neo model on my 24gb 3090 and can fit a 30b parameter OPT model when combining my 3090 and 12gb 3080
- mdda 4y agoCould you point to any resources online about how to do this? e.g. is this using 8-bit quantisation?