9 ms·
Neurogenesis Deep Learning
- groar 10y agoBasically trying to achieve a certain level of plasticity in deep neural nets by getting inspiration from https://en.wikipedia.org/wiki/Adult_neurogenesis https://en.wikipedia.org/wiki/Adult_neurogenesis
- andreyk 10y agoTo add on to this - they "specifically consider the case of adding new nodes to pre-train a stacked deep autoencoder", by basically keeping track of when certain layers cannot reproduce their input and then adding more nodes+retraining with both new (not reproduced) and old data. It is quite intuitive, basically the most naive and obvious first attempt at the problem (not meant in a condescending way, just want to point out it's not that generalizable and is pretty ad-hoc).
- jostmey 10y agoNeurogensis? How about neural death as a way to prune large neural networks into more compact ones--now that is a research idea!
- 10b5-1 10y agoLooks like these researchers are trying to make a network more adaptive, I think that deleting nodes would only make them worse at the current task they're being trained on as well as worse on the tasks they're being adapted to. You could train a model using neurogenesis to increase its accuracy, and then use distillation to train a smaller network to comparable accuracy. But these are two very different, but complementary, problems.
- rdlecler1 10y agoYou're assuming that all nodes are functionally important/non-spurious.
- 10b5-1 10y agoI'm not assuming that, I'm giving the model more options and letting it decide what is functionally important/non-spurious. It might take it longer, but I don't assume that.
- leegao 10y agoMore parameters also means that the likelihood of overfitting (the training set) increases. Currently (and rather unintuitively, considering that ML is an applied optimization field, and optimization is usually concerned with underfitting), the bane of ML is overfitting. It's easy to supply a model with high representational capacity, but it's impossible to learn anything interesting in a reasonable amount of time. You'll learn how to fit your training set perfectly because your model has enough degrees of freedom to let you fit a million points arbitrarily well, but that doesn't mean that the resulting fit describes the data in a meaningful way. This is why a core tenet of ML is to prune parameters whenever possible. Neurogenesis increases representational capacity whenever it detects that your underlying model does not have sufficient representational capacity to fit the data; from this perspective, you start small (undercapacity) and then you gradually increase your capacity until you hit the optimal model. In other words, Neurogenesis is also a way for you to minimize the number of options. On the other hand, giving the model with more options than it necessarily needs and letting it decide what is important will usually backfire. Rather than learning a few meaningful/functional features, it can just go ahead and completely fit the training data from the very beginning. It will therefore decide that everything is important, because all those extraneous parameters will let it squeeze that last 0.5% out of your training set.
- antome 10y agoThis idea is at least partially in use with regularisation and dropout. The difference at least with dropout is that the "killed" neurons are then massaged back into the network in order become useful again.
- shmageggy 10y agoAgreed that this is another way of framing the problem of regularizing a network. Rather than starting with a big network and penalizing complexity, they are starting with a simple network and adding complexity. To that end, I'd've liked to see a comparison to dropout or L1/L2 regularization.
- visarga 10y agoBiological neurons themselves are stochastic so they have an internal "dropout" that doesn't seem to hurt, on the contrary, these perturbations and imperfect communication increase learning ability.
- cing 10y agoI believe several papers have examined efficiently pruning neural networks, but neural death would be better branding ;) (https://arxiv.org/pdf/1506.02626v3.pdf https://arxiv.org/pdf/1506.02626v3.pdf https://arxiv.org/pdf/1510.00149v5.pdf https://arxiv.org/pdf/1510.00149v5.pdf)
- simonster 10y agoYann LeCun called a technique for pruning weights "optimal brain damage" back in 1990: http://yann.lecun.com/exdb/publis/pdf/lecun-90b.pdf http://yann.lecun.com/exdb/publis/pdf/lecun-90b.pdf
- SimonKinds 10y agoThere is also this one https://arxiv.org/abs/1506.02515 https://arxiv.org/abs/1506.02515 which takes pruning a step further to reduce the sparsity. Also this one https://arxiv.org/abs/1608.04493 https://arxiv.org/abs/1608.04493 which makes sure not to kill any neurons which proves to be useful at a later stage in the pruning process.
- kajecounterhack 10y agohttps://arxiv.org/abs/1503.02531 https://arxiv.org/abs/1503.02531 Modern applications of small networks regularly reduce sizes from larger state-of-the-art networks using distillation. Distillation compacts neural networks while affecting accuracy minimally. Instead of pruning directly from the large network, just learn how it generalizes. Takes fewer nodes / overall operations (Multiplications / Additions).
- Vitrified 10y agoNow that is interesting. I hadn't realized methods to combine trained networks so efficiently were already readily available.
- kajecounterhack 10y agoCertain companies use these methods to make state of the art neural nets work on your phones :) Also "combine" might not be the right word, since it's really transfer learning. "Distill" is really a descriptive verb. Maybe my original wording was confusing; I shouldn't have said "distillation compacts" -- distillation is a process by which you can create a more compact version of a complex neural net.
- hyperbovine 10y agoSo basically, give your deep networks drugs and alcohol.
- chriswarbo 10y agoAs others have mentioned, there are approaches like regularisation and dropout which try to do similar things. What I find interesting is the fact there are two reasons to do this: to generalise/avoid-overfitting and to reduce resource usage. It seems like almost all effort is spent on the former, since everyone's aiming for higher accuracy numbers. Are there any widely-used methods to tackle the latter? For example, I'm imagining a system which is either given measurements of its resource usage (time, memory, etc.) or uses some simple predictive model (e.g. time ~ number of layers * some constant), and works within some resource bound: - If we're below the bound, expand the model (add neurons, etc.) to allow accuracy increases (note "allow": it's ok to ignore/regularise-to-zero the extra parameters to avoid overfitting) - If we're above the bound, prune the model (in a way which tries to preserve accuracy) - Allocate resources to optimise some objective, e.g. reduce variance by pruning the parameters of the best-performing class/predictor/etc. and using those resources to expand the worst performer. The closest thing I know of are artificial economies, but they seem to be more like a selection mechanism (akin to genetic programming) than a direct optimisation procedure (like gradient descent on an ANN).
- visarga 10y agoThere are many ways to compress networks - by pruning neurons, by enforcing sparsity, by representing activations and gradients on one bit (or a few bits), and by transfer learning where a large net is transferred into a smaller one.
- chriswarbo 10y agoYes, my question was more about meta-level algorithms for balancing size against performance. Especially adaptive methods such that we're not just growing up to a limit and stopping, but selectively allocating resources to those parts which need them. Adapting over time would be nice too: "thinking harder" when there are idle resources, but shrinking the results back down under load.
- SimonKinds 10y agoThis paper http://dl.acm.org/citation.cfm?id=2830854 http://dl.acm.org/citation.cfm?id=2830854 kind of has a solution to being more efficient. It has two networks and uses the smaller one (more efficient) to infere first. If the result is accurate with high probability (the probability of one class is much larger than the probability of any other class) then there is no need to run the big (expensive) network.
- irinarish 10y agoExactly! For that, check out our OpenReview ICLR submission on NEUROGENESIS-INSPIRED DICTIONARY LEARNING: ONLINE MODEL ADAPTION IN A CHANGING WORLD, by Sahil Garg, Irina Rish, Guillermo Cecchi, Aurelie Lozano https://openreview.net/revisions?id=HyecJGP5ge https://openreview.net/revisions?id=HyecJGP5ge
- gallerdude 10y agoVery wishful thinking on my part, but I think we're far closer to a general intelligence than most expect.
- empath75 10y agoI think what we might see is a kind of autonomous corporation that is nominally under the control of shareholders, a CEO or a board, but which makes decisions without very much or any human input, and which gains some amount of legal rights through corporate personhood. It won't be a 'general ai', though. More like a set of loosely connected systems that operate 'in the best interests of the shareholders', however that's defined. It's pretty much the end state of the trend of pushing decision making to algorithms to remove moral and legal culpability from individuals.
- BeingIncubated 10y agoI'm hoping that eventually translates to the state.
- hmate9 10y agoOut of curiosity: how far away do you think we are?
- gallerdude 10y agoI'll say... 5 - 7 years. This is all based on pure speculation, and being a little more than a ML hobbiest. One thing is for sure though - when we do reach that point, everything changes forever.
- hmate9 10y agoSlightly off topic, but I hate how publications are written. It seems like authors are purposely using big words and sentences that are often 5-6 lines long in order to make it seem more clever. I find myself often having to reread a sentence in order to understand it. These algorithms are often very simple and can be easily explained. Don't over complicate them.
- amelius 10y agoThen here's a challenge: could you write the abstract of the article in "simple English", without changing the meaning?
- unfamiliar 10y agoHere's my attempt; not a huge number of changes because it was not too bad to begin with, but with slightly less self-indulgent language, and a lot of the jargon has to stay (partly because I don't know the field): Neural machine learning methods, such as deep neural networks (DNN), have achieved remarkable success in a number of complex data processing tasks. These methods have arguably had their strongest impact on tasks such as image and audio processing - areas where humans have always performed better than conventional algorithms. In contrast to biological neural systems, which are capable of learning continuously, deep artificial networks have a limited ability for incorporating new information after a network has been trained. As a result, continuous learning methods could be very helpful in allowing deep networks to handle data sets which change over time. Here, inspired by the process of adult neurogenesis in the hippocampus, we investigate how adding new neurons to artificial neural networks can allow them to acquire new information, while preserving what they have already learned. Our results on the MNIST handwritten digit dataset and the NIST SD 19 dataset, which includes lower and upper case letters and digits, show that neurogenesis looks like a good approach for tackling the "stability-plasticity dilemma" that has been a problem for adaptive machine learning algorithms for some time. As an academic, I tend to agree that we frequently feel compelled to apply more verbosity than is strictly required in order to communicate the intended semantic constructs.
- 10y ago
- joantune 10y agoIt never ceases to amaze me that the best steps towards achieving AI is to look at how we perceive that a Neuron works and simulate it. And the thing is, we aren't exactly sure why exactly that is.. it's amazing. Sometimes the best thing we can do is imitate nature
- spott 10y agoThis isn't strictly true though. Spiking Neural Networks [0] attempt to be more accurate representations of human neurons, but haven't really caught on because they aren't really much better than our perceptron model of neurons, at least for the things we are trying to do with them. [0]http://www.ane.pl/pdf/7146.pdf http://www.ane.pl/pdf/7146.pdf
- joantune 10y agonice, thanks for sharing, interesting read so far (read the introduction), will definitely give it a better look out of curiosity
- spynxic 10y agoI find that somewhat strange. Why attribute the idea of introducing new nodes to a graph to biological concepts? It seems like a simple step in exploration, similar to how one might think to vary the weights of the nodes randomly over some range.. unless there is some technique biology uses to pre-configure the nodes upon introduction to the network, that might be rather interesting.
- joantune 10y agoBecause they tried to model neurogenesis, the same way that artificial neural networks were invented while trying to mimic some parts of how neurons work I guess..
- deleted 10y ago[deleted]
- partycoder 10y ago
- paulsutter 10y agoTL/DR: - "We specifically consider the case of...a stacked deep autoencoder (AE), which is a type of neural network designed to encode a set of data samples such that they can be decoded to produce data sample reconstructions with minimal error - "The first step of the NDL algorithm occurs when a set of new data points fail to be appropriately reconstructed by the trained network...When a data sample’s RE is too high, the assumption is that the AE level under examination does not contain a rich enough set of features to accurately reconstruct the sample. - "The second step of the NDL algorithm is adding and training a new node, which occurs when a critical number of input data samples (outliers) fail to achieve adequate representation at some level of the network. - "The final step of the NDL algorithm is intended to stabilize the network’s previous representations in the presence of newly added nodes. It involves training all the nodes in a level with both new data and replayed samples from previously seen classes on which the network has been trained.
- iverjo 10y agoHow does this relate to Progressive Neural Networks [0]? That technique is also about accumulating knowledge (while not forgetting existing knowledge) [0] https://arxiv.org/abs/1606.04671 https://arxiv.org/abs/1606.04671
- argonaut 10y agoSorry if I'm being snobbish, but I do wonder why this paper is only being submitted to IJCNN, a 2nd tier machine learning conference. I know students who publish undergrad research at workshops with lower acceptance rates than IJCNN. I can't think of any important machine learning papers published in IJCNN in the recent past.
- habitue 10y agoIt depends on what conclusions you're trying to draw from that information. What conference a paper was accepted to is a second-order signal of the noteworthiness. It's probably easier for someone versed in the field to just read the paper to determine if it's interesting. If you're using the conference as a quick pass/fail as you skim through the abstracts of hundreds of papers, ok, but you probably wouldn't make time to comment on HN about it in that case. This paper looks like it builds on pretty well-known techniques like stacked autoencoders, so let's see what first-order noteworthiness data we can gather from a quick skim of the paper. If I had to guess why it wasn't accepted into a better conference: - It uses stacked autoencoders, which are pretty out of fashion - It bothers reporting results on MNIST - (more subjectively) It pulls an unfortunately common technique of saying "here's something the brain does" and then hand-waving that it's a deep reason why a technique they've come up with is useful, when in fact the relationship is just "inspired by the general idea of", not "performs the same function as" the biological mechanism. In this case, I think the tenuous connection of their technique to research on neurogenesis is pretty flimsy. Clearly neurogenesis is not how an adult human brain forms new memories or gains proficiency in new skills (which they acknowledge in the conclusion)
- argonaut 10y ago> It's probably easier for someone versed in the field to just read the paper to determine if it's interesting. > If you're using the conference as a quick pass/fail as you skim through the abstracts of hundreds of papers, ok You answered your own statement, I think. Most researchers will skip a paper in a second tier conference. In fact, most I know won't read an entire paper - they'll only read some of it and skip stuff. You're correct that I am not an active researcher (otherwise I would not have time to be commenting). I merely did some research back in college. But honestly that little experience gives me a huge leg up on most HN commenters in understanding research. It's unfortunate that the only reason this paper is #1 on HN is because it has a cool title. That being said, MNIST is not really a disqualifier. (Unfortunately) MNIST is the most popular dataset referenced in NIPS 2016 papers (https://twitter.com/benhamner/status/805864969065689088 https://twitter.com/benhamner/status/805864969065689088). The handwaving is also forgivable; many NIPS papers handwave a lot too.
- m3kw9 10y agoVery nice coin term for just another deep learning method
- SubiculumCode 10y agoThey simulate neurogenesis, I guess, but they do not incorporate the most interesting part of that neurogenesis: That is the new neurons are born into the dentate gyrus, a region thought to have a particular capacity to orthoganalize feature representations that are similar (e.g. pattern separate) allowing distinct memories to be formed for similar events. The dentate gyrus outputs to a region called Cornu ammonis 3 (CA3) which is heavily recurrent, and thought to br able to pattern complete a full representation from partial inputs. That is, CA3 can encode and retrieve the relations between 2 or more features or objects. For a mathematical model and review one might read: Rolls (2013) https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3812781/ https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3812781/ but many others exist. I'd write more but typing in my phone is driving me to distraction.
- arkymark 10y agothis is really interesting, where/how did you learn this? I'd like to learn more about these things - brain regions, connections, functions - and what they might imply about the kinds of computations that are going on, but my background is mainly on the AI/math side of things.
- SubiculumCode 10y agoMy dissertational research was on the development of the subfields of the hippocampus in childhood. So these papers from the rat literature were relevant and often inspiring.
- SubiculumCode 10y agoI'd like to add that our knowledge of the details of hippocampal neuroanatomy are probably the most advanced of any brain region, which allows the somewhat informed construction of computational models. I wish I had more time to learn modeling methods, I have specific developmental hypotheses I'd like to test in such a model. In the end, I'd probably need to find a knowledgable collaborator though.
- infinite8s 10y agoOne place to start is the "Principles of Neural Science", the intro textbook to neuroscience.
- yahyaheee 10y agoSeems a bit like a GAN
- irinarish 10y agoFor a model that incorporates both neuronal birth and death, see ICLR submission at OpenReview: https://openreview.net/revisions?id=HyecJGP5ge https://openreview.net/revisions?id=HyecJGP5ge NEUROGENESIS-INSPIRED DICTIONARY LEARNING: ONLINE MODEL ADAPTION IN A CHANGING WORLD Sahil Garg, Irina Rish, Guillermo Cecchi, Aurelie Lozano
- irinarish 10y agoA good point was made that a model of neurogenesis must also incorporate neuronal death besides neuronal birth (since hippocampus and the brain as a whole have physical constraints, you can't keep growing your network infinitely :). That's why any model of neurogenesis must incorporate interplay between birth and death of new (and old) neurons; that's was the main idea of the paper I mentioned in an earlier post (this year ICLR submission https://openreview.net/forum?id=HyecJGP5ge https://openreview.net/forum?id=HyecJGP5ge) Note that just adding nodes to networks was proposed before, eg. the classical work on cascade correlations.