3 ms·
For inference, the best models are so large they won't fit in System RAM. GPT@home is not going to make a difference in that scenario. For training such large
by Vetch 4y ago
For inference, the best models are so large they won't fit in System RAM. GPT@home is not going to make a difference in that scenario.
For training such large models, data parallelism is no longer sufficient and tensor/pipeline parallelism is required. The problem is communication bottlenecks, differing device/network speeds and massive data transfer requirements become serious enough issues to kill any naive distributed training across the internet approach. Deep learning companies use fancy 100Gbps+ connections, do kernel hacking and use homogeneous hardware and it's still a serious challenge. There is no incentive for them to invest in something like GPT@home.
But it's not impossible and there's some research being done in the area. Although, it'll be a while until a GPT@home approach becomes a ready alternative. See https://arxiv.org/abs/2206.01288 https://arxiv.org/abs/2206.01288 and their recent GPT-JT test for more. Another development would be for networks to become more modular.
- nullc 4y agoRam isn't terribly expensive, it's not unreasonable to have 1 or 2 TB of ram. 1TB costs about $3500 as 64GB dimms. (some of my 4u hosts have 96 ddr4 sockets too... though 6tb of ram is getting a little pricey. :)) > use fancy 100Gbps+ connections, you can pick up 100gbps mellanox nics on ebay for $50 on a good day, $200 whenever. If you're only connecting up two or three hosts you can just use multiport cards and a couple dac cables, rather than a switch. I suspect for inference though there is a substantial locality gain if you're able to batch a lot of users into a single operation, since you can stream the weights through while applying them to a bunch of queries at once. But that isn't necessarily lost on a single user, it would be nice to see a dozen distinct completions at once.