5 ms·
NoProp: Training neural networks without back-propagation or forward-propagation
- gwern 1y agohttps://www.reddit.com/r/MachineLearning/comments/1jsft3c/r_noprop_training_neural_networks_without/ https://www.reddit.com/r/MachineLearning/comments/1jsft3c/r_... I'm still not quite sure how to think of this. Maybe as being like unrolling a diffusion model, the equivalent of BPTT for RNNs?
- cttet 1y agoIn all their experiments, backprop is used for most of their parameter though...
- hansvm 1y agoThere is a meaningful distinction. They only use backprop one layer at a time, requiring additional space proportional to that layer. Full backprop requires additional space proportional to the whole network. It's also a bit interesting as an experimental result, since the core idea didn't require backprop. Being an implementation detail, you could theoretically swap in other layer types or solvers.
- ActorNightly 1y agoI think we need to start thinking about one shot training. I.e instead of context into LLM, you should be able to tell it a fact, and it will encode that fact into the updated weights.
- itsthecourier 1y ago"Years of works of the genetic algorithms community came to the conclusion that if you can compute a gradient then you should use it in a way or another. If you go for toy experiments you can brute force the optimization. Is it efficient, hell no."
- isentrop1c 1y agoAre you just a bot stealing reddit comments? (https://www.reddit.com/r/MachineLearning/comments/1jsft3c/r_noprop_training_neural_networks_without/mlnsehp/ https://www.reddit.com/r/MachineLearning/comments/1jsft3c/r_...)
- itsthecourier 1y ago"Whenever these kind of papers come out I skim it looking for where they actually do backprop. Check the pseudo code of their algorithms. "Update using gradient based optimizations""
- arrakark 1y agoSame. If I had to guess it's just local gradients, not an end-to-end gradient.
- f_devd 1y agoI mean the only claim is no propagation, you always need a gradient of sorts to update parameters. Unless you just stumble upon the desired parameters. Even genetic algorithms effectively has gradients which are obfuscated through random projections.
- erikerikson 1y agoNo you don't. See Hebbian learning (neurons that fire together wire together). Bonus: it is one of the biologically plausible options. Maybe you have a way of seeing it differently so that this looks like a gradient? Gradient keys my brain into a desired outcome expressed as an expectation function.
- red75prime 1y ago> See Hebbian learning The one that is not used, because it's inherently unstable? Learning using locally accessible information is an interesting approach, but it needs to be more complex than "fire together, wire together". And then you might have propagation of information that allows to approximate gradients locally.
- erikerikson 1y agoIs that what they're teaching now? Originally it was not used because it was believed it couldn't learn XOR (it can [just not as perceptrons were defined]). Is there anyone in particular whose work focuses on this that you know of?
- sriku 1y agoPosted on 31st March which would've been 1st April somewhere else in the world?
- erikerikson 1y agoWe have gradient free algorithms: Hebbian learning. Since 1949?
- uoaei 1y agoAnd there's good reasons why we use gradients today.
- sva_ 1y agoThat's more a theory/principle, not an algorithm by itself.
- erikerikson 1y agoIt is an update rule: Wij = f(Wij, xi, xj) The weight of the connection between nodes i and j is modified by a function over the activations or inputs of node i and j. The are many variants of back propagation too. Regardless, yes it would be used within a network model such as a Hopfield network.
- erikerikson 1y agoSee also https://en.m.wikipedia.org/w/index.php?title=Generalized_Hebbian_algorithm&wprov=rarw1 https://en.m.wikipedia.org/w/index.php?title=Generalized_Heb...
- hansvm 1y agoIt's a neat idea. It's not too dissimilar in spirit from gradient boosting. The point about credit assignment is crucial, and that's the same reason most architectures and initialization methods are so convoluted nowadays. I don't really like one of their premises and conclusions: > that does not learn hierarchical representations There's an implicit bias here that (a) traditional networks do learn hierarchial representations, (b) that's bad, and (c) this training method does not learn those. However, [a] is situational, and it's easy to construct datasets where a standard gradient-descent neural net will learn a different way, even with a reverse hierarchy. [b] is unproven and also doesn't make a lot of intuitive sense to me. [c], even in this paper where they make that claim, has no evidence and also doesn't seem likely to be true.
- DrFalkyn 1y agoIf we could ever figure out what wet brains actually do (continuous feedback, ? enzyme release ? ) this might be possible
- tsimionescu 1y agoKeep in mind that our brains also have a great deal of built in trained structure from evolution. So even if we understood exactly how a brain learns, we may still not be able to replicate it if we can't figure out the highly optimized initial state from which it starts in a fetus.
- gpjt 1y agoPresumably that is limited by the gig or so of information in our DNA, though?
- tsimionescu 1y agoThe amount of information transmitted from one generation to the next is potentially much more than the contents of DNA. DNA is not an encoding of every detail of a living body, it is a set of instructions for a living body to create an approximate copy of itself. You can't use DNA, as far as we know, to create a new organism from scratch to create a new organism without having the parent organism around to build it. We do know for certain that many parts of a cell divide separately from the nucleus and have no relation to the DNA of the cell - most well known being the mitochondria, which have their own DNA, but also many organelles just split off and migrate to the new cell quasi-independently. And this is just the simplest layer in some of the simplest organisms - we have no idea whatsoever how much other information is transmitted from the parent organism to the child in ways other than DNA. In particular in mammals, we have no idea how actively the mother's body helps shape the child. Of course, there's no direct neuron to neuron contact, but that doesn't mean that the mother's body can't contribute to aspects of even the fetal brain development in other ways.
- gpjt 1y ago
- toxik 1y agoAn interesting idea for sure, but why only evaluate it on 28x28 pixel images? Why is their flow matching so much worse in some cases? Missing some analysis. Their words on it say nothing: > For CIFAR-100 with one-hot embeddings, NoProp-FM fails to learn effectively, resulting in very slow accuracy improvement In general any actual analysis is made impossible because of the lack of signal in the results. Fig 5 tells me nothing when the span is 99.58 to 99.46 percent accuracy.