4 ms·
In what way does it not parallelize well? There are mounds of research in federated learning.
by question_away 7y ago
In what way does it not parallelize well? There are mounds of research in federated learning.
- kyle_grove 7y agoIn fact, one of the chief advantages of the BERT/Transformer architecture over ELMO/LSTM is the ability to parallelize.
- bitL 7y agoRNNs (LSTM/GRU) tend to have issues with scaling. Attention-based models like Transformer on the other hand scale extremely well.
- RocketSyntax 7y agoI've read that you can't split up large layers to be trained on separate processors either horizontally (one layer per processor) or vertically (parts of many layers).
- pheug 7y agoOn a shared memory system there's little need to do that - there's much more parallelism to be had from accelerating fine grained operations, like matrix multiplications to compute each layer's output. On a distributed system, splitting up layers between machines to do distributed training is pretty much what Google initially designed Tensorflow for. Generally it scales less well due to the need to communicate massive amounts of data between nodes and much lower network throughput than what GPU/TPU memory provides.