5 ms·
Not a deep learning expert, but: it seems that without backpropagation for model updates, the communication costs should be lower. And that will enable models t
by BenoitP 4y ago
Not a deep learning expert, but: it seems that without backpropagation for model updates, the communication costs should be lower. And that will enable models that are easier to parallelize?
Nvidia isn't creating new versions of its NVLink/NVSwitch products just for the sake of it, better communication must be a key enabler.
Can someone with deeper knowledge can comment on this? Is communication a bottleneck, and will this algorithm uncover a new design space for NNs?
- jerpint 4y agoIt’s more that without backpropagation, you no longer need to store your forward activations across many layers to compute the backwards pass which usually is dependent on a forward pass. When a network is hundreds of layers, and batches are very large, the forwards and backwards accumulations add up in terms of memory required. Communication across GPUs doesn’t solve this but instead allows to have either many models running in parallel on different GPUs to increase Barch size or to share many layers across GPUs to increase model size. Quick communication is critical to maintain training speeds that aren’t astronomical
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- PartiallyTyped 4y ago> will this algorithm uncover a new design space for NNs? No. Hinton "discovered" stacking ensembles and gave it a new name, fancy analogies to biological brains and then made it worse. The gist of this is that you can select a computational unit, be it a linear layer, or a collection of layers, compute the derivative of the output with respect to the parameters, and update them. Each computational unit is independent, meaning that you don't calculate gradients going outside of it. This is the same as training a bunch of networks, computing predictions, and then using another layer to combine the predictions. This is called stacking, and the networks are called an "ensemble". You can do this multiple times and have N levels of meta estimators. Instead of fitting the ensemble and then the meta estimator, Hinton proposes training both simultaneously but without allowing gradients to flow through. That is stupid because if you don't allow gradients to flow through, you will see a context drift as the data distribution changes. Hinton observed this context drift, to deal with that, he proposed normalizing the data. On one extreme, you can use individual linear units as the models, and on the other extreme, you can combine all units into a single neural network and treat that as a module. So no, this does not open any new design, it's an old idea, worsened, and wrapped in fancy words and post-facto reasoning. If you are curious how a linear layer is an ensemble, observe that each vector is its own linear estimator, making the linear mapping an ensemble of estimators.
- sdwr 4y agoThats a lot less cheaty, biologically speaking, than full backprop. This Hinton guy sounds like he knows what he's talking about. "Context drift as data distribution changes" sounds a hell of a lot like real life to me. Normalized = hedonic treadmill on long view At the micro scale, data that overflows the normalization is stored in emotional state, creating an orthogonal source of truth that makes up for the lack of full connected learning.
- chamwislothe2nd 4y agoYou're accusing one of the foundations of modern AI with either being a fraud, or incompetent. At best that seems short sighted, no?
- nayroclade 4y agoYou're obviously new to Hacker News :-D
- chamwislothe2nd 4y agoUnfortunately, I've been here for a decade in one form or another. Every now and then someone writes something so pompous that I just can't help myself but post. Back to lurking now. Cheers!
- PartiallyTyped 4y agoIf we can't publicly scrutinize people who have great sway in the industry, what does that say about us as a research community? The fact that I argued why I found it bogus based on well established principles, and I get shitted on by people who by all means have provided nothing to this conversation and except suppressing criticism or throwing ad-hominems should tell all about the quality of discourse. Dismissing criticism, not by arguments, but by the mere name of the person does a disservice to everyone. If the research can't stand on its own, independent of the author, then it is not good research.