20 ms·
How deep is the brain? The shallow brain hypothesis
- jakobson14 3y agoIf I had a nickel for every time some neurologist tried to compare brains to neural networks. It's a surefire way to tell someone is either desperate for grant money or has been smoking crack. (previously: comparing brains and "electronic computers") Their entire article hinges on the complaint "brain seems shallow and neural networks are deep, ergo neural networks are doing it wrong." Neurologists seem to have a really hard time comprehending that researchers working on neural networks aren't as clueless about computers as neurology is about the brain. They also vastly overestimate how much engineers working on neural networks even care about how biological brains work. Virtually every attempt at making neural networks mimic biological neurons has been a miserable failure. Neural networks, despite their name, don't work anything like biological neurons and their development is guided by a combination of A) practical experimentation and refinement, and B) real, actual understanding about how they work. The concept of resnets didn't come from biology. It came from observations about the flow of gradients between nodes in the computational graph. The concept of CNNs didn't come from biology, it came from old knowledge of convolutional filters. The current form and function of neural networks is grounded in repeated practical experimentation, not an attempt to mimic the slabs of meat that we place on pedestals. Neural networks are deep because it turns out hierarchical feature detectors work really well, and it doesn't really matter if the brain doesn't do things that way. And then you have the nitwits searching the brain for transformer networks. Might as well look for mercury delay line memory while you're at it. Quantum entanglement too.
- deleted 3y ago[deleted]
- junofan 3y agoI dunno, failure seems okay. Wouldn’t expect a better paradigm to beat SOTA at first. It’s totally plausible that neurons use eg. transposons in a way we don’t yet have the instrument resolution to characterize, which would suggest that you don’t need 1000 layers, but a lookup table or something.
- b33j0r 3y agoAgreed, but I do also think that order emerged from chaos. It’s an easy claim when order is defined by itself! But in reality, we’re equipped exactly to exist, and we still wonder why in a backwards way, even with education (guilty!) AI is the task of playing God like toddlers at recess, and LLMs the tower of babel. I still wanna play, it’s fun
- robbrown451 3y agoI can't agree with the dismissiveness of this comment, and frankly I find its tone out of line and not with the spirit of Hacker News. There are insights that can come from studying the brain, that do indeed apply. Some researchers may not glean anything from such studies, and some may. I have no doubt that as neural networks get more an more powerful, we will continue to find more ways they are similar to the brain, and apply things we've learned about the brain to them. I certainly prefer to see people making comparisons of neural networks to the brain, that the old "it's just a glorified autocomplete" and the like. Relax.
- krainboltgreene 3y agoWhat does this comment add to the discussion?
- robbrown451 3y agoI dunno. My comment complained about the parent comment not adding positively to the discussion. And gave at least a bit of support for that complaint. Would you have preferred I emulate your style, and complain while providing no support for my complaint? Ok.
- krainboltgreene 3y agoBeing positive is not a requirement of commenting on HN, but you should comment with something that is substantive, so yes I do think you shouldn't have commented at all. Tone policing is cringe.
- robbrown451 3y agoExactly what are you doing here then? But hey I guess I can do this too. How's this? Using cringe as an adjective is cringe.
- krainboltgreene 3y ago
- mrstone 3y agoA neurologist is a medical doctor. Neuroscientists are the PhDs who do the actual research.
- crustacean111 3y agoCNNs actually are biologically inspired. The receptive field in a CNN mimics the way that cortical neurons only respond to stimuli in a restricted region of the visual field. Different cortical neurons have receptive fields that partially overlap to cover the whole visual field [1]. [1] - https://en.wikipedia.org/wiki/Convolutional_neural_network https://en.wikipedia.org/wiki/Convolutional_neural_network
- jakobson14 3y agoYou're going to have to dig deeper. The concept of a receptive field goes all the way back to convolutional filters. It's not surprising that we found out later the brain also uses such a fundamental element of signal theory.
- SubiculumCode 3y agoOh good. So you do admit that there are useful parallels between signal processing, statistical processing, and the brain.
- vkou 3y agoSure, and airplanes are inspired by birds. That doesn't mean that detailed studies of the Boeing 747 are going to unlock a lot of hitherto unknown mysteries of heron behaviour.
- jacobsimon 3y agoI mean, I know you’re just providing an analogy, but people are still studying the physics of bird flight and we’re nowhere close to building machines yet that can maneuver the way birds can. https://www.quantamagazine.org/geometric-analysis-reveals-how-birds-mastered-flight-20220803/ https://www.quantamagazine.org/geometric-analysis-reveals-ho...
- ben_w 3y agoI could believe "we have more to learn", but not "we're nowhere close": https://youtu.be/w6VLzKACnS8?si=DZgOPuBRG4Vt98su https://youtu.be/w6VLzKACnS8?si=DZgOPuBRG4Vt98su
- jacobsimon 3y agoThis is a really weird take. There is such a long history of shared insights between biology and neural network research, and to say they’re unrelated or can’t take inspiration from one another is bizarre. > The concept of CNNs didn't come from biology I just opened a survey paper on CNNs and literally the first sentence of the paper reads: > “Convolutional Neural Network (CNN) is a well-known deep learning architecture inspired by the natural visual perception mechanism of the living creatures. In 1959, Hubel & Wiesel [1] found that cells in animal visual cortex are responsible for detecting light in receptive fields. Inspired by this discovery…” Source: https://arxiv.org/pdf/1512.07108.pdf%C3%A3%E2%82%AC%E2%80%9A https://arxiv.org/pdf/1512.07108.pdf%C3%A3%E2%82%AC%E2%80%9A
- jakobson14 3y agoThat's later backfill, a retroactive change to give a manufactured "biological" origin story. Whether they're real or not, researchers love a good "we took this from nature, isn't nature wonderful!" explanation. The C in CNN isn't "Convolution" for no reason. It came from work with convolutional filters (yay Sobel kernels!) which at it's height became filter banks and gabor filters and so on before neural networks pretty much killed off handcrafted feature development. Every explanation of how CNNs work still falls back to the original convolutional kernel intuition.
- dartos 3y agoYou can use that argument for anything you disagree with. Do you have a source or anything?
- jakobson14 3y agoHave a read through the first paper describing a convolutional neural network, from 1998: http://yann.lecun.com/exdb/publis/pdf/lecun-01a.pdf http://yann.lecun.com/exdb/publis/pdf/lecun-01a.pdf There's absolutely no mention of biological inspiration whatsoever. At the same time, one can point to a long and rich history of convolutional filters being used in signal processing. And then there's the name, Convolutional Neural Network. The entire concept of a CNN is framed as a series of learned filters.
- two_in_one 3y agoWhile I agree with this emotional post there is one nuance. Neural networks aren't intelligent, brain is. And that's where we want to be. Checking gradients and studying filters can get us only this far. So, using brain as inspiration looks like a good option. There are other, but nobody knows where next breakthrough will be. Like nobody knew five years back that transformers are so powerful. My guess next step to AGI will be a complex modular multi-modal system. With hierarchy, workers and controllers, complex signals.. Sound familiar? Brain is sort of it. This is need for embodied AI, obviously. But, interesting thing, it's needed even for body-less AGI too. I.e. AGI is not a big calculator (!), it's more like real-time system. One reason is that full search is impossible. So, in many cases requests will be like 'give the best answer you can find in 4 seconds'. 'and keep looking'. So far we have only real-time dumb robots and NN big calculators. And brains, of course.
- dilawar 3y ago> previously: comparing brains and "electronic computers") Before that: comparing brain with hydraulic machines. There has been tendency to compare brain with most complex machine known to us at that particular time. "Descartes was impressed by the hydraulic figures in the royal gardens, and developed a hydraulic theory of the action of the brain. We have since had telephone theories, electrical field theories, and now theories based on computing machines… . We are more likely to find out how the brain works by studying the brain itself, and the phenomenon of behavior, than by indulging in far-fetched physical analogies." -- Karl Lashley 1951
- spindle 3y agoAnd also comparing brains to clockwork.
- mjburgess 3y agoI cannot agree enough with Karl here. What is the brain? An organic system with deep roots in the organic body, with deep causal connections with its environment. There's little sense in ignoring the whole basic mode of operation, physics, chemistry and biology of the brain in order to analogise it to another system without any of those properties. This, at best, provides a set of inspirations for engineers -- it does nothing for science.
- ben_w 3y agoI mildly disagree (although your final conclusion is correct: it indeed does nothing for science). The deepest fundamental structures in the brain[0] are quantum fields, which are also the deepest fundamental structures in everything else. There is no known quantum field of "soul" or "intelligence". The right abstraction is higher, and could still be a whole lot of things; but as maths can be implemented in logic, which can be implemented in electronics or clockwork or hydraulics, it doesn't matter what analogy is used — and my mild disagreement here is that such inspiration has been useful and gotten us this far. [0] that we know of
- mjburgess 3y ago
- wslh 3y agoOnly an observer of the topic but I think it is good to review Koch's book about the real complexity of a single neuron [1]. [1] https://www.amazon.com/Biophysics-Computation-Information-Computational-Neuroscience/dp/0195181999/ref=sr_1_1?keywords=Biophysics+of+Computation%3A+Information+Processing+in+Single+Neurons&qid=1561229482&s=books&sr=1-1 https://www.amazon.com/Biophysics-Computation-Information-Co...
- SubiculumCode 3y agoIf you read this article, I think most would understand that it is primarily aimed at other neuroscientists, and only using ML structures an an analogy only, and I think a somewhat useful one to boot. The real point of the article was to propose a general hierarchy for how information flows in the brain, to emphasize the importance of subcortical brain even in higher order cognition, and proposes how simultaneous processing of multiple levels of representation can inform action and thought. As a developmental neuroscientist, I found the article insightful and thought provoking. Further, it is quite consistent with major hypotheses in psychology, how the hippocampus works (a subcortical structure) and combines information into memories: See fuzzy trace theory [1], for example. Your dismissive tone is unappreciated, ill-informed, and crass. [1] https://en.wikipedia.org/wiki/Fuzzy-trace_theory https://en.wikipedia.org/wiki/Fuzzy-trace_theory
- visitor4711 3y agofully agree
- bjourne 3y agoFirst, I wonder how you got access to the article? It is behind a paywall and not yet uploaded to the sites I usually find paywalled articles on. Second, there is no need to compare brains to neural networks because brains are neural networks. Neurons form vertices and axons edges connecting the aforementioned. What you are perhaps thinking of are artificial neural networks - most of which are very dissimilar to brains. But even then you are wrong. Artificial Izhikevich and Hodgkin-Huxley neural networks attempts to closely mimic the behavior of real neurons. While deep, hierarchical artificial neural networks have been more successful than biologically plausible ones, that may be because the technology isn't ready yet. After all, the perceptron was invented in the 1950's but didn't become prominent until the 2010's (or so). Perhaps we need new memories that better map to (real) neural network topologies, or perhaps 3d chips that can pack transistors in the same way brains pack neurons.
- mjan22640 3y agoA neuron is analogous to a 3d integrated circuit rather to a transistor. A molecule acts like a transistor https://medium.com/the-physics-arxiv-blog/the-origin-of-life-and-the-hidden-role-of-quantum-criticality-ca4707924552 https://medium.com/the-physics-arxiv-blog/the-origin-of-life... Changes in mechanical pressure, electric field, other molecules attachment, photon absorption, can control the conductivity. Organic semiconductors designed to fit like lego bricks to naturally build the desired structure are IMHO the way to go to produce 3d circuits, rather than layered silicone litography.
- ben_w 3y ago> silicone I've seen this particular mistake a lot recently. New and exciting auto-corrupt from the latest version of iOS? Given that our brains rewire themselves live, which ANNs can only do by being excessively connected and updating weights to/from zero, silicone (I'm thinking mainly the oil form) may be a better inspiration than lego. https://en.wikipedia.org/wiki/Silicone https://en.wikipedia.org/wiki/Silicone
- mjan22640 3y ago
- __loam 3y agoAs a biomedical engineer who went into software, thank you for this comment lol. So tired of rehashing this.
- andromaton 3y agoBooks and articles I was reading in the 80s (eg Minsky and Papert, Byte magazine) were referring to Rosenblatt and retinas.
- blovescoffee 3y agoDude. What holy and special work do you do? There's nothing dumb or dull in searching for analogous structure between two effective machines, neither of which we understand.
- nathias 3y agoMetaphores and analogies are important tools of thinking, even in science, some bear fruits some lead to errors, but we can't know in advance.
- hliyan 3y ago"brain seems shallow and neural networks are deep, ergo neural networks are doing it wrong" Please don't claim things the author didn't. What I read was "ergo (artificial) neural networks may be missing a trick"
- deleted 3y ago[deleted]
- radarsat1 3y ago> every time some neurologist tried to compare brains to neural networks Value of this comment aside, it kind of makes me chuckle how casually it (and other comments in this thread) just drops the word "artificial" from neural networks here, specifically when comparing with neurology. The irony is funny. Like, somehow we've forgotten why we call them that in the first place, exactly when talking about the thing that inspired the approach.
- svara 3y agoDoesn't know what a neurologist is, knows they do shit work.
- NoToP 3y agoI disagree profoundly. There are things the brain does we have not yet been able to reproduce with a neural network, or to the extent we have seemingly with excessive resources of training and network size. Therefore there is some salient feature of neurology which has been overlooked. I don't think it is necessary to mimic biology down to the exact function of real neurons, but there must in fact be something we are neglecting to mimic.
- ben_w 3y agoPossibly, but it may also be that we're training them wrong. "Book smart, not street smart" (to use a catchphrase) would apply perfectly to GPT models: brain the size of a rodent's, with 50,000 year's experience of reading Reddit, Wikipedia, and StackOverflow, but no "real life" experiences of its own.
- sheeshkebab 3y agoIt’s indeed odd that current dnn’s require massive amount of energy to retrain and lack any kind of practical continuous adaptation and learning.
- quickthrower2 3y agoWith computer-based intelligence we have the overhead of computing every bit though (probably) inefficient silicon and direct electric currents. The brain leverages the properties of chemicals, though millions of years of evolution.
- jakobson14 3y agoThe brain isn't a faster computer. An infinitely-fast computer wouldn't meaningfully change the "expensive training vs fast, static inference" workflow that neural networks have always been developed around (except in the most brute force-y "retrain on the entire world, every single nanosecond" sense).
- quickthrower2 3y agoI think we agree? I am talking to the efficiency of the brain. Not processing speed. Efficiency of the brain to do things advantageous to the selfish genes I guess. The brain is supremely efficient at what the brain has evolved to do. It is almost tautological! Because if it wasn't, it wouldn't have evolved to that. Silicon comes from an alien land, and is emulating. Even with the best algorithms there has to be a limit on how efficient a computer-based intelligence can be without changing how the chips work. You could spin it around and say, well computers are better at many things than humans, and there is no way you could get a biological brain to be as good for the same amount of power (e.g. a raspberry pi can do calculations our brain couldn't possibly do).
- DiggyJohnson 3y agoReally well said, I think this is an excellent way to frame the dichotomy (comparison?). Much of these threads make the binary mistake: can these systems be compared, or are they fundamentally different? A bit of both, almost certainly.
- Salgat 3y agoThe brain communicates with itself, so deep layers are equivalent to sections of the brain talking to each other. The only relevance white matter depth has is with regard to how it's trained, and since it doesn't use gradient descent, it's irrelevant to neural networks in that regard.
- blovescoffee 3y agoIntercommunication does not equal layer depth.
- Salgat 3y agoWhy not? All a deep neural network is doing is progressive data transformations into something more abstract and meaningful to later layers.
- beaugunderson 3y agohttps://anonymfile.com/dR8a/s41583-023-00756-z.pdf https://anonymfile.com/dR8a/s41583-023-00756-z.pdf
- hliyan 3y ago"brain seems shallow and neural networks are deep, ergo neural networks are doing it wrong" Please don't claim things the author didn't. What I read was "ergo (artificial) neural networks may be missing a trick"
- hliyan 3y agoIgnore. Reposted this under correct parent comment
- phlogisticfugu 3y agodeep learning models have already been permitting "shallow signals" for a while. see "skip connections" https://theaisummer.com/skip-connections/ https://theaisummer.com/skip-connections/
- low_tech_punk 3y agoReplay of Jeff Hawkins group’s A Thousand Brains theory?
- SubiculumCode 3y ago"his theory" lol. Jeff Hawkins is a bit player
- rando_dfad 3y agoOriginal to Jeff or not, "A Thousand Brains" does a decent job presenting an interesting and highly plausible model of how the neocortext may function. Your comment would be very valuable to me if it included pointers to better sources. I have sufficient background to see gaps in Jeff's book, and would be interested in exploring these, perhaps through the references you seem to be aware of.
- MagicMoonlight 3y agoIf it was shallow then it wouldn’t take 25 years for a human brain to fully train. The fact that some parts of it need that much data mean they must be way up the hierarchy.
- GranularRecipe 3y agoThe reason for deep learning is that shallow networks are very hard (or impossible) to train. In that sense, long time of training is evidence for shallow networks.
- IshKebab 3y agoNo it's because shallow networks can't express complex functions. If you think about it the shallowest network is pretty much a lookup table. They can theoretically model any function, but the number of parameters needed means in practice they can't. Deep networks can learn much more complex functions for the same number of parameters.
- GranularRecipe 3y agoWhat the ratio for the number of parameters required to learn some complex function between a shallow network and a deep network (preferably as a function of the complexity)?
- PartiallyTyped 3y agoI mean… a 3 layer network is a Universal approximator… and you can very much do network distillation… it’s just that getting them wide enough to learn whatever we want them to isn’t computationally efficient. You end up with much larger matmuls which let’s say for simplicity exhibit cubic scaling in the dim. In contrast, you can stack layers and that comes much more computationally friendly because your matmuls are smaller. Of course you then need to compensate with residuals, initialisation, normalisation, and all that, but it’s a small price to pay for scaling much much better with compute.
- 3y ago
- Simon_ORourke 3y agoJudging by some of the levels of driving around these parts, the brain may be very shallow indeed.
- deleted 3y ago[deleted]
- bjornsing 3y ago> This shallow architecture exploits the computational capacity of cortical microcircuits and thalamo-cortical loops that are not included in typical hierarchical deep learning and predictive coding networks. As I understand it the thalamus is basically a giant switchboard though. I see no reason to believe that it never connects the output of one cortical area to the input of another, thus doubling the effective depth of the neural network. (I haven’t read this paper though, as it was behind a paywall.)
- audunw 3y agoI seem to remember research stating that an individual neuron has very complex behaviour that requires several ML “neurons” / nodes to simulate. So if you do a comparison, perhaps the brain is deeper than you’d think by just looking at the graph of neurons and their synapses. Could we construct a neutral net from nodes with more complex behaviour? Probably, but in computing we’ve generally found that it’s best to build up a system from simple building blocks. So what if it takes many ML nodes to simulate a neuron? That’s probably an efficient way to do it. Especially in the early phase where we’re not quite sure which architecture is the best. It’s easier to experiment with various neural net architectures when the building blocks are simple.
- magicalhippo 3y ago> Could we construct a neutral net from nodes with more complex behaviour? Well there's spiking neural networks (SNN)[1], which are modeled more closely to how neurons actually work. Main obstacle is still, as far as I know, that there's no way to train a SSN as efficiently as a "regular" neural network, which lends itself very nicely to gradient descent and similar[2]. [1]: https://en.wikipedia.org/wiki/Spiking_neural_network https://en.wikipedia.org/wiki/Spiking_neural_network [2]: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9313413/ https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9313413/
- rmorey 3y ago> I seem to remember research stating that an individual neuron has very complex behaviour that requires several ML “neurons” / nodes to simulate. This is probably what you're remembering: https://www.sciencedirect.com/science/article/pii/S0896627321005018 https://www.sciencedirect.com/science/article/pii/S089662732...
- rsrsrs86 3y agoThe brain backprops??????
- andbberger 3y agothere is no evidence to support this
- chriskanan 3y agoThe brain has a lot of skip connections and is massively recurrent. In a sense, the brain can be thought of as having infinite depth due to recurrent thalamno-cortical loops. They do mention thalamno-cortical loops in the paper, so I think a more concrete definition of what is meant by "depth" would be helpful.
- lausbub 3y agoThe "infinite depth" seems to be a matter of definition. It's practically infinite if you include feedback loops via learning. If you exclude learning, then it's far from "infinite". Activations linger for up to 15-30 seconds, so at oscillations of around 30 Hz that would result in about 450-900 loops (times an unknown small multiplier for the actual number of layers). But the brain presumably only backprops/optimizes a few layers at a time and not much "through" time.
- sudosysgen 3y agoThere's also evidence that the brain does optimize through time and might be implementing, at least in some places, algorithms close to LSTD.
- rsrsrs86 3y agoBeyond the mere topological metaphor of neural networks there is almost nothing in common between brains and widigital computation. This is a widespread fallacy of category.
- epgui 3y agoI completely disagree, and I think this is an example of human-exceptionalism bias.
- rando_dfad 3y agoand more specifically, between chemical-based information processing systems and Von Neumann architectures for binary information processing. Agreed, a widespread fallacy of category. But computers still do some pretty cool things. Powerful tools.
- naasking 3y ago> Beyond the mere topological metaphor of neural networks there is almost nothing in common between brains and widigital computation. I mean, sure, but the topology is exactly what makes both work, so we only really care about the topology.
- lawrenceyan 3y agoWe have skip connections and recurrent neural networks at home.
- lawlessone 3y agoSo does this mean DNN are in some ways deeper than human brains?
- spacetimeuser5 3y agoWho finally cares how exactly an ANN matches a human brain? Is such ANN smarter than ChatGPT? It is more useful to use AI to develop more ecologically valid measurement methods for biology.