4 ms·
Couldn’t you still copy by training a new network on a new device to have same outputs for the same inputs as the original?
by dsabanin 2y ago
Couldn’t you still copy by training a new network on a new device to have same outputs for the same inputs as the original?
- tomxor 2y agoYes, but training is the most expensive part of ML, for example GPT-3 is estimated to cost something like 1-4 million USD. With ANN you can do it one time and then clone the result for negligible energy cost. Maybe training a batch of PNNs in parallel could save some of the energy cost, but I don't know how feasible that is considering they could behave slightly differently during training causing divergence... Now that sarcastic comment at the bottom of this thread is starting to sound relevant "Schools".
- kmmlng 2y ago> Yes, but training is the most expensive part of ML, for example GPT-3 is estimated to cost something like 1-4 million USD. That entirely depends on how many inferences the model will perform during its lifecycle. You can find different estimates for the energy consumption of ChatGPT, but they range from something like 500-1000 MWh a day. Assuming an electricity price of $0.165 per kWh, that would put you at roughly $80,000 to a $160,000 a day. Even at the lower end of $80,000 a day, you'll reach your $4 Million in just 50 days.
- tomxor 2y agoThat's not a proportional comparison, n simultaneous users to 1 training. How many users across how many GPUs is that 80k? With PNN you would have to multiply n by 1-4 million, training cost explodes.
- l33tman 2y agoThat's not true for the most well-known models. For example Meta's LLAMA training and architecture was predicated on the observation that training cost is a drop in the well compared to the inference cost for a model's lifetime.
- etiam 2y agoDistillation (as you may be aware). https://arxiv.org/abs/1503.02531 https://arxiv.org/abs/1503.02531 Having to do that in each instance is still really cumbersome for cheap mass deployment compared to just making a digital-style exact copy, but then again I guess a main argument for wanting these systems is that they'd be doing things unachievable in practice on digital computers. In some cases one might be able to distill to digital arithmetic after the heavy parts of the optimization are done, for replication, distribution, better access for software analysis, etc.