5 ms·
Backprop Alternative: Augmented Lagrangian Predictive Coding
https://arxiv.org/abs/2605.31022 https://arxiv.org/abs/2605.31022
https://github.com/SakanaAI/pc-alm https://github.com/SakanaAI/pc-alm
- guld 18d agoNew paper by Sakana.ai [1] [1]: https://arxiv.org/abs/2605.31022 https://arxiv.org/abs/2605.31022
- cs702 18d ago~85% accuracy on MNIST. Sigh. How does it do on CIFAR-10, or even better, ImageNet? Interesting research, not sure it's a backprop alternative. === EDIT: accuracy on MNIST is not ~90%. It's ~85%.
- bz_bz_bz 18d agoTheir image classification benchmarks include both: https://pub.sakana.ai/pc-alm/assets/figures/benchmark_accuracy_depths.webp https://pub.sakana.ai/pc-alm/assets/figures/benchmark_accura...
- cs702 18d ago~74% on CIFAR-10. Still a far cry from backprop. I didn't see ImageNet. TinyImageNet is something else.
- rocketships 17d agoThere might be only one paper that's trained on full Imagenet using methods like these. Training a Predictive Coding Network on ImageNet using Equilibrium Propagation Tugdual Kerjan, Rasmus Høier, Benjamin Scellier https://arxiv.org/abs/2606.03584 https://arxiv.org/abs/2606.03584 It's quite an undertaking.
- Lerc 18d agoIt might be beneficial while not being optimal on its own. The obvious example is if it has different behaviour around local minima, it could be an altenate pathway out. I have often wondered if doing training with radically different aproaches for the first few iterarions would avoid any method specific artifacts before the weights had time to denoise.
- qarl 18d agoThey state replacing backprop is not their goal. Their goal is to understand how distributed systems which cannot do backprop (the brain) can still do learning.
- im3w1l 18d agoPersonally I think the dirty secret of the brain is that a lot of things are hard coded. And many things that we need to learn are also hard coded except that some parameters need to be tuned. If we puke, the brain will not do general aversive learning, it will learn to avoid specifically the last thing eaten, because it instinctively knows about food poisoning. Imprinting is absolutely fascinating. Some newborn animals will run a very simple pattern detector like looking for a red dot or something and use that to bootstrap their conception of their parent. For fully general learning I have a hunch that it can be done using local history plus a semi-global reward scalar (global neurotransmittor levels).
- DoctorOetker 17d agoregardless if the intelligence in the brain is hardcoded or not, to the extent it is, this information must have been compressed in the genome, which runs counter to almost all observations: a child doesn't remember the experience of their ancestors, for example. The only sense in which we do carry mental state without relearning is emotions, instincts, reflexes (some neuronal pathways that connect the eye to the middle ear), hormonal driven behavior (fear adrenalin). For another, there are about 200k promotor regions (including non-coding) in the human genome. A promotor region might have say 6 to 15 bits of information. Can you compress 2025 or even 2024 era LLM intelligence into 3 megabit = ~400 kB ? I think not. I think a lot of compression is still possible, but 400 kB? So I think we can box up the idea of "dirty secrets of the braing: not learning but hard coding". There is a lot of hard coding in biology, but brains are evolved specifically to enable learning within the individual lifetime instead of only learning by natural selection. I also don't buy the following argument: > If we puke, the brain will not do general aversive learning, it will learn to avoid specifically the last thing eaten, because it instinctively knows about food poisoning. Each time it happens that I end up puking, I do feel aversion and try to avoid puking at all, sometimes I succeed but sometimes is just puke. There must be fundamental puke reflexes (which one fails to avoid) and avertable puke reflexes.
- imtringued 17d agoIt's really sad that they are only a few years away from making backpropagation completely obsolete.
- AIorNot 18d agoOh wow the theoretical implications in neuroscience exite me here - is this a potential model of Fristons Markov Blanket concept “ Probably the most ambitious and all-encompassing version of the ‘Bayesian turn’ in cognitive science is the free energy principle (FEP). The FEP is a mathematical framework, developed by Karl Friston and colleagues (Friston, Kilner, and Harrison 2006; Friston et al. 2010; Friston 2010; Friston et al. 2017a; Friston 2019), which specifies an objective function that any self-organizing system needs to minimize in order to ensure adaptive exchanges with its environment. One major appeal of the FEP is that it aims for (and seems to deliver) an unprecedented integration of the life sciences (including psychology, neuroscience, and theoretical biology). The difference between the FEP and earlier inferential theories (e.g., Gregory 1980, Grossberg 1980, Rao and Ballard 1999, Lee and Mumford 2003) is that not only perceptual processes, but also other cognitive functions such as learning, attention, and action planning can be subsumed under one single principle: the minimization of free energy through the process of active inference (Friston 2010; Friston et al. 2017). ”
- nullbio 17d agoIsn't the FEP basically just loss minimization over KL-divergence? In other words, it's the same thing we already do with ML and already have been doing for years? I've never understood where this differs to the status quo, or why this isn't just a relabelling of techniques/concepts. Although I didn't look too deeply.
- AIorNot 17d agoYeah you are right to basics of the paper and I am extrapolating a bit here: I think the remarkable result of this paper is that they add a local Lagrange multiplier λ at each layer, which accumulates constraint/prediction error over the inference dynamics. At equilibrium, in the linear case, those local multipliers converge to exactly the same gradient signal that backpropagation would calculate globally Now what is Predictive coding: its a network that can minimize prediction errors through local recurrent interactions instead of an explicit global backward pass. Now I am making the leap to Fristons more philosophical and mathematical work not the paper - so that is me making the allusion But a light bulb moment for me dawned when I read it This process (PC-ALM) gives us a concrete example of how globally coherent inference/credit assignment can emerge from purely local dynamical interactions. PC-ALM lets a recurrent dynamical system relax toward a state in which the backprop gradient is represented locally throughout the network. That distinction is potentially important for neuroscience. A brain doesn’t obviously have a central routine saying: loss.backward() it certainly has recurrent neural populations whose states continuously influence neighboring populations. This paper is demonstrating that, at least mathematically, those sorts of local recurrent dynamics can generate the same credit information that backprop obtains through the chain rule. The authors explicitly motivate predictive coding as a biologically plausible local-learning alternative because standard BP requires globally coordinated error variables and update ordering. Think about it also give plausible evolutionary to chain intelligence through cells coming together and creating nested networks This has got to be how the neurological intelligence sausage gets made What it eventually means for ML I’m Not sure but hopeful it opens a door
- lukeinator42 18d agoThere is a lot of interesting research into predictive coding as an alternative means to solve the credit assignment problem that might be a more plausible model of what happens in the brain. I really liked this paper that showed using a predictive coding learning rule leads to the exact same gradients as backprop in arbitrary networks: Predictive Coding Approximates Backprop Along Arbitrary Computation Graphs https://direct.mit.edu/neco/article/34/6/1329/110646/Predictive-Coding-Approximates-Backprop-Along https://direct.mit.edu/neco/article/34/6/1329/110646/Predict...
- rocketships 17d agoThis paper uses the "fixed prediction assumption" so I think has caused some confusion (i.e. it's not PC, but PC with a small bandaid). It's a great paper though just the title is misleading somewhat. But all of Beren Millidge's papers are quite good and he was so prolific in the space. Cool to have seen him on the Dwarkesh podcast recently as well.
- ACCount39 17d agoThat was genuinely a strong signal that whatever learning algorithm the brain uses isn't "magic", and probably can be approximated with the ML tools we have. It also pointed at the possibility that the learning algorithms brain uses might be, like the paper has demonstrated, less compute-optimal and data-optimal than backprop - but far easier to implement in a localized, distributed fashion. Because backprop requires activation storage and global connectivity - and both "hot" data storage and connectivity of any kind are extremely expensive for the brain.
- mikelitoris 18d agoRolls right off the tongue
- Jeff_Brown 18d agoCould this relate to continual learning? It lets you update without pausing the entire system.
- DoctorOetker 17d agoNothing prevents interleaving inference and gradient descent / RMAD / backpropagation from being interleaved. Is a man a father or a son? It's a false dilemma, it can be both. Is a training step pretraining or finetuning? It's the same mathematical operation, with the only difference the intention or reason of applying the training step.
- verdverm 17d agoLoRA throws a tiny amount of sand in these gears, but I generally agree there is near zero difference beyond the semantics we mere mortals assign to written tokens
- red75prime 17d agoMaybe it has some interesting properties that could prevent catastrophic forgetting, but so far it looks like a biologically plausible approximation of backpropagation and it might be useful to decrease the training compute requirements.
- mike_hearn 17d agoYou can do that with backprop too. Nothing says you can't inference on a set of weights at the same time as you produce an updated copy from them. They say in the paper what it's about: mostly just scientific curiousity but such approaches might be useful for making DNNs more energy efficient via neuromorphic hardware in future. For continual learning at the weight level there's the business model issue. The labs are already deep in the red, the last thing they want is to give up shared weights. The I/O and storage costs of that would make it infeasible. Already KV caches are a sort of dynamic 'fast weights' and those are expensive!
- Tuna-Fish 17d ago
- rao-v 17d agoI wonder if you could take a traditional backprop trained LLM and apply this approach to finetuning it (presumably needs less memory and compute?). It could be another entry in the spectrum between LORA and full fine tuning.
- rao-v 17d agoTurns out this is more compute intensive than regular backprop. The win (if any) would be in training on future highly distributed architectures where neighbouring parts of the model have accesss to very high bandwidth (to each other) but communicating to further away parts is more expensive
- txhwind 17d agoNice introduction to a simple but useful idea! The Lagrangian works like a time-smoothed optimizing direction state, but it can be placed on any wire, even at non-differentiable boundary! Can it be better than existing training methods for discrete components like argmax, MoE or VQ-VAE? Maybe networks can be composed by a lot of learnable discrete components, or even bits and gates finally.
- txhwind 17d agoother questions: - can it be used to relax timing order requirement in pipeline parallelism? Each node update lambda on communication, and optimize weights at other time. - given that BP is using SGD, can batches and T share the same timeline in optimization, while keeping the descent direction?
- deleted 17d ago[deleted]
- erichocean 17d ago> PC-ALM closes the PC-BP gap at matched inference budget in nonlinear networks Okay, so no efficiency improvement?
- Tuna-Fish 17d agoBackprop has significant costs, if a local solution truly matches its performance, it will be dropped like a hot potato.
- erichocean 17d agoAgreed, but how will we find out? Pre-training alone on frontier-class models costs billions of dollars.
- anon291 17d agoThe brain doesn't need to do back propagation. Back propagation is a simulation of gradient descent which the universe achieved naturally by its own devices. Indeed if you tie the loss function to some sort of excited energy state in such a way then leave it to be, the universe will naturally drive the system to its local minima.
- redwood 17d agoReminds of https://www.neurofounders.co/articles/bezos-backed-flourish-raises-500-million-for-brain-inspired-ai https://www.neurofounders.co/articles/bezos-backed-flourish-...