5 ms·
As long as the cost for training a model is in the 7+ figures, that means that any such open model is bound to be tracked to someone with deep enough pockets to
by ayende 3y ago
As long as the cost for training a model is in the 7+ figures, that means that any such open model is bound to be tracked to someone with deep enough pockets to sue.
Consider that you just spend a few millions on training a model on copyrighted data. Release it would reveal that, problem.
I _guess_ you can try doing training in the public, like Seti @ Home or something like that, which distributes the risk? But no idea if this is even possible in this context.
- vlovich123 3y agoIs the training cost equally that high if you do adversarial training?
- 127361 3y agoBut if we have 100,000 people with RTX 4090s mining some new AI based cryptocurrency, that happens to train the model in the process, it's going to be a highly effective system. And we can do it anonymously over Tor or I2P.
- PeterisP 3y agoYou really can't, because bandwidth matters as much as compute power - you can only utilize as much power as you can transfer data to/from. The training methods for current LLMs are parallellizable only by very frequently transferring all the data back and forth, and needing a node to gather and merge all the updates very frequently, and redistribute it to every other node so that they can make any progress. And a GPU that's not connected you with a high-speed link is pretty much useless as you can't make useful progress until you get their part pack, and "their part" is very large (i.e. the update size you need to get back is comparable to all of the model size) and you need to do that very frequently. Training on nVidia many-GPU pods works because of high-speed interconnect (e.g. 600 gigabytes per second for 8 GPUS in A100 pod), and if your internet bidirectional speed is much less than 600gbps, then if you have 100000 free remote RTX 4090s, you simply can do the compute locally faster than you can exchange information with the other GPUs.
- deleted 3y ago[deleted]
- 127361 3y ago"This paper presents a distributed model-parallel training framework that enables training large neural networks on small CPU clusters with low Internet bandwidth." Low bandwidth being <1Gbps. They've also tested it with GPUs as well. https://arxiv.org/pdf/2201.12667 https://arxiv.org/pdf/2201.12667 Maybe there's the possibility of a completely new AI architecture that can still be efficiently trained when there are very low bandwidth connections between nodes? Specifically targeting this use case would make sense, given all the millions of underutilized GPUs out there in peoples' desktop computers. Also https://arxiv.org/pdf/2106.10207 https://arxiv.org/pdf/2106.10207 ?
- spacebanana7 3y agoThe cost of modifying open-ish source models with copyrighted data is much lower. For example, the cost of modifying Mixtral with uncensored data is currently about $1200 [1] [1] https://youtu.be/GyllRd2E6fg?si=SJmPLsPlCRRT0uPV&t=236 https://youtu.be/GyllRd2E6fg?si=SJmPLsPlCRRT0uPV&t=236
- wongarsu 3y ago> I _guess_ you can try doing training in the public, like Seti @ Home or something like that, which distributes the risk? But no idea if this is even possible in this context. Let's say that's a field of active research. Right now you need very low latency, to the point that people connect training clusters via infiniband instead of ethernet despite the computers being meters apart. But approaches that tolerate internet-level latency are being developed
- throwuwu 3y agoI’m willing to bet that there are projects at every chip company looking into how to drive down these costs. The first step is to put the architecture of the model and the backprop into silicon. The weights are the only variables so if you can reduce those to cheap and fast modules that can be mass produced you could come up with something cheaper than a GPU. If the profitability of running and training this specific model is high enough than its worth the investment of time and money to set up the production line.
- Der_Einzige 3y agoThe costs are falling 10-100x every few months. If you told people in 2022 that they could run 70b paramater models at 2 bit quantization on your home 4090, they'd have laughed in your face. They're not laughing anymore. SFT/DPO are ultra efficient compared to RLHF, and innovation in this space is only just now getting started.