7 ms·
For everyone reading neither the article nor the paper: - both show neural networks can learn the game of life just fine - the finding is that to learn the ru
by moconnor 2y ago
For everyone reading neither the article nor the paper:
- both show neural networks can learn the game of life just fine
- the finding is that to learn the rules reliably the networks need to be very over-parameterised (e.g. many times larger than the minimal size needed for hand-crafted weights to perfectly solve the problem)
This is not really a new result nor a surprising one, nor does it say anything about the kinds of functions a neural network can represent.
It's an attempt to understand an existing observation: once we have trained a large overparameterized neural network we can often compress it to a smaller one with very little loss. So why can't we learn the smaller one directly?
One of the theories referred to in the article and paper is the lottery hypothesis, which states that a large network is a superposition of many small networks and the larger you are the more likely at least one of those gets a "lucky" set of weights and converges quickly to the right solution. There is already interesting evidence for this.
- actionfromafar 2y ago> once we have trained a large overparameterized neural network we can often compress it to a smaller one with very little loss. So why can't we learn the smaller one directly? I feel something similar goes on in us humans. Interesting to think about.
- rmnclmnt 2y agoIndeed, you have to over-engineer something before converge to a leaner solution to a problem
- brookst 2y agoAnd I think this is because the ideal complex system is one where all the subsystems and parts combine to produce adequate reliability. Exponentiation means it is more efficient to start by far exceeding the required reliability and then optimizing the most expensive subsystems/parts. It is less efficient and far more frustrating if multiple things have to be improved to meet requirements.
- actionfromafar 2y agoThis is really insightful.
- nopinsight 2y agoThere is indeed an analogous process in the brain. "The number of synapses in the brain reaches its peak around ages 2-3, with about 15,000 synapses per neuron. As adolescents, the brain undergoes synaptic pruning. In adulthood, the brain stabilizes at around 7,500 synapses per neuron, roughly half the peak in early childhood. This figure can vary based on individual experiences and learning." -- written by GPT-4o Confirmed by e.g. https://extension.umaine.edu/publications/4356e/ https://extension.umaine.edu/publications/4356e/
- patcon 2y agoAlso perhaps why the evolved pattern of death is important: a subnetwork is selected in a brain, which is suited to a specific geological, physical, biological and cognitive environment that the brain is navigating. But when the environment shifts beneath the organisim (as culture does and the living world in general does), then the subnetwork is no longer the correct one, and needs to be reinitialized. Or in other words, even in an information theoretic sense, it's true: you can't teach a old dogs new tricks. You need a new dog.
- yumong 2y agoFrom an evolutionary perspective, wouldn't we expect to have developed a way to "reset" parts of our brain then?
- sitkack 2y agoThat is what kids are for.
- actionfromafar 2y agoHaha, exactly. "We did".
- ianmcgowan 2y agoThis reminds me of a hacker news comment that blew my mind - basically "I" am really my genetic code, and this particular body "I" am in is just another computer that the code has been moved to, because the old one is scheduled to be decommissioned. So I am really just the latest instance of a program that has been running continuously since the first DNA/RNA molecules started to replicate.
- interroboink 2y agoIf you're interested in such things, then start layering on epigenetics. The "I" is a product not just of genes, but of your environment as you developed. I was just reading about bees' "royal jelly" recently, and how genetically identical larvae can become a queen or a worker based on their exposure to it. So the program is not just the zeroes and ones, so to speak, but also more nebulous real-time activity, passed on through time. Like a wave on the ocean.
- deleted 2y ago[deleted]
- seanhunter 2y agoYes. My naive intuition about this is you need the extra parameters precisely to do the learning because learning a thing is more complicated than doing the thing once you have learned how. There are lots of natural examples that fit this intuition eg in my mind "junk" DNA is needed because the evolutionary mechanism is learning the sequences which work in a similar way. You don't need all that extra DNA once you have it working but once you have it working there's little selection pressure to clean up/optimise the DNA sequence so the junk stays.
- mysecretaccount 2y agoThanks, very clear explanation!
- G3rn0ti 2y ago> the lottery hypothesis Isn’t that another way of saying the optimization algorithm used in finding the network‘s weights (gradient descent) can not find the global optimum? I mean this is nothing new, the curse of dimension prevents any numeric optimizer to completely minimize any complicated error function and it’s been known for decades. AFAIK there is no algorithm that can find the global minimum of any function. And this is what currently limits neural network models: They could be much simpler and less resource hungry if we had better optimizers.
- staunton 2y agoIn practice, you don't want the global optimum because you can't put all possible inputs in the training data and need your system to "generalize" instead. Global optimum would mean overfitting.
- WithinReason 2y agoRegularisation should not be done with the optimiser but with the loss function and the architecture.
- redox99 2y agoRegularization only helps you so much.
- staunton 2y agoMaybe it should not be done but the large neutral networks this decade absolutely rely on this. A network at the global minimum of any of the (regularized) loss functions that are used these days would be waaay overfitted.
- uoaei 2y agoThe entire reason SGD works is because the stochastic nature of updates on minibatches is an implicit regularizer. This one perspective built the foundations for all of modern machine learning. I completely agree that the most effective regularization is inductive bias in the architecture. But bang for buck, given all the memory/compute savings it accomplishes, SGD is the exemplar of implicit regularization techniques.
- DrScientist 2y agoThis may be a silly question but - so rather than train a big network and hope a subnetwork wins the lottery - why not just train a smaller network with multiple runs with different starting weights?
- MalphasWats 2y agoLargely for the same reason people don't "just win" the lottery by buying every possible ticket (any more).
- scarmig 2y agoThe larger network contains exponentially more subnetworks. 10x the size contains far more than 10x subnetworks (although it'd also take more than 10x as long to train).
- DrScientist 2y agoAh ok - so you are saying the explosion of subnetworks is higher than the explosion of training time - leading to a win. In this case a positive benefit of combinatorial complexity.
- tomxor 2y ago> So why can't we learn the smaller one directly? The lottery hypothesis intuitively makes sense, but as an outsider I find this concept for evaluating learning methods really interesting - To hand craft a tiny optimal networks for simple yet computationally irreducible problems like GoL as a way to benchmark learning algorithms. Or is it more than that? for a sufficiently small network maybe there aren't that many combinations of "correct solutions", so perhaps the way the network emerges internally could really be interrogated by comparison.
- dilyevsky 2y agoIsn’t this basically the idea behind dropout technique?
- Grimblewald 2y agoNo,the idea behind dropout is to reduce an over-reliance on specific outputs thereby, in theory and typically in practice, making the network learn more reliable representations reducing the chance of overfitting.