4 ms·
I have a probably really stupid question but I'm going to ask it because I really don't know the answer. What's stopping someone from starting a distributed ef
by brhsagain 4y ago
I have a probably really stupid question but I'm going to ask it because I really don't know the answer.
What's stopping someone from starting a distributed effort to train an open source ChatGPT model, like Folding@home but for AI? Is it a technical limitation with how models are trained, difficulty of coordination, something else?
- code-is-code 4y agoTraining in parallel on more then one GPU is pretty hard. Distributing the training over many random computers and adding at least some kind of validation, seems to kill the performance. So it would always be easier to rend cloud GPUs.
- coffeebeqn 4y agoBut would it be cheaper?
- zamadatix 4y agoThe models are extremely large to ship around, sync, and load on consumer hardware over the internet. Maybe there are better ways to build and train a model for a distributed use case though.
- zirgs 4y agoThe model that's used by Stable Diffusion is only around ~4 GB in size.
- KeplerBoy 4y agowhich would make traditional distributed training prohibitively slow. You would have to sync the weights with residential internet speeds (10s of MB/s) instead of PCI-e Speeds (10s of GB/s).
- zirgs 4y agoIf you want to add new stuff to Stable Diffusion - you don't have to retrain the main model from scratch. You can train it only on a few hundred images at a time and then add the resulting model as an extension or merge it into the main model. People train their models on a single celebrity or a single artist that way. Other types of AI could be trained in a similar way.
- zamadatix 4y agoIs this amenable to merging from thousands in one go or do you have to train -> merge -> train -> merge to not overwrite each other's trainings?
- yucky 4y agoSounds like we have our use-case for blockchain!
- deleted 4y ago[deleted]
- wendyshu 4y agoHard to split the problem into small enough pieces
- igxtickckg 4y agoworking on it. together.xyz
- shaklee3 4y agobesides the answers already given, the network is already the bottleneck on supercomputers on a local network. training that over a WAN would exacerbate that issue and make training even slower.
- imtringued 4y agoThat is actually irrelevant. I couldn't try out flan-t5-xl but the model flan-t5-model is like an inferior GPT but it is opensource. I don't necessarily think that the size of the model is that significant. What makes ChatGPT so impressive is that you don't have to mess around with the settings vs Flan T5. More effort has to be put into obtaining good training data than just throwing compute at the problem.