8 ms·
It seems the take home is weight decay induces sparsity which helps learn the "true" representation rather than an overfit one. It's interesting the human brain
by greenflag 3y ago
It seems the take home is weight decay induces sparsity which helps learn the "true" representation rather than an overfit one. It's interesting the human brain has a comparable mechanism prevalent in development [1]. I would love to know from someone in the field if this was the inspiration for weight decay (or presumably just the more equivalent nn pruning [2]).
[1] https://en.wikipedia.org/wiki/Synaptic_pruning https://en.wikipedia.org/wiki/Synaptic_pruning
[2] https://en.wikipedia.org/wiki/Pruning_(artificial_neural_network) https://en.wikipedia.org/wiki/Pruning_(artificial_neural_net...
- visarga 3y agoThe inspiration for weight decay was to reduce the capacity to memorize of the model until it perfectly fits the complexity of the task, not more not less. A model more complex than the task is over-fitting, the other one is under-fitting. Got to balance them out. But the best cure for over-fitting is to make the dataset larger and ensure data diversity. LLMs have datasets so large they usually train one epoch.
- nightski 3y agoIt sounds nice in theory, but the data itself could be problematic. There is no temporal nature to it. You can have duplicate data points, many data points that are closely related but describe the same thing/event/etc.. So while only showing the model each data point once ensures you do not introduce any extra weight on a data point, if the dataset itself is skewed it doesn't help you at all. Just by trying to make the dataset diverse you could skew things to not reflect reality. I just don't think enough attention has been paid to the data, and too much the model. But I could be very wrong. There is a natural temporality to the data humans receive. You can't relive the same moment twice. That said, human intelligence is on a scale too and may be affected in the same way.
- visarga 3y ago> I just don't think enough attention has been paid to the data, and too much the model. I wholly agree. Everyone is blinded by models - GPT4 this, LLaMA2 that - but the real source of the smarts is in the dataset. Why would any model, no matter how its architecture is tweaked, learn about the same ability from the same data? Why would humans be all able to learn the same skills when every brain is quite different. It was the data, not the model And since we are exhausting all the available quality text online we need to start engineering new data with LLMs and validation systems. AIs need to introspect more into their training sets, not just train to reproduce them, but analyse, summarise and comment on them. We reflect on our information, AIs should do more reflection before learning. More fundamentally, how are AIs going to evolve past human level unless they make their own data or they collect data from external systems?
- ben_w 3y ago> It was the data, not the model It's both. It's clearly impossible to learn how to translate Linear A into modern English using only content written in pure Japanese that never references either. Yet also, none of the algorithms before Transformers were able to first ingest the web, then answer a random natural language question in any domain — closest was Google etc. matching on indexed keywords. > how are AIs going to evolve past human level unless they make their own data? Who says they can't make their own data? Both a priori (by development of "new" mathematical and logical tautological deductions), and a posteriori by devising, and observing the results of, various experiments. Same as us, really.
- riversflow 3y agoI see this brought up consistently on the topic of AI take-off/X-risk. How does an AI language model devise an experiment and observe the results? The language model is only trained on what’s already known, I’m extremely incredulous that this language model technique can actually reason a genuinely novel hypothesis. A LLM is a series of weights sitting in the ram of GPU cluster, it’s really just a fancy prediction function. It doesn’t have the sort of biological imperatives (a result of being complete independent beings) or entropy that drive living systems. Moreover, if we consider how it works for humans, people have to _think_ about problems. Do we even have a model or even an idea about what “thinking” is? Meanwhile science is a looping process that mostly requires a physical element(testing/verification) to it. So unless we make some radical breakthroughs in general purpose robotics, as well as overcome the thinking problem I don’t see how AI can do some sort tech breakout/runaway.
- crdrost 3y agoAnd there have been a lot of approaches to do this, my favorite one being the idea that maybe if we just randomly zap out some of the neurons while we train the rest, that forcing it to acquire that redundancy might privilege structured representations over memorization. Just always seemed like some fraternity prank, “if you REALLY know the tenets of Delta Mu Beta you can recite them when drunk after we spin you around in a circle twelve times fast!”
- whimsicalism 3y agohttps://nitter.net/Yampeleg/status/1688441683946377216 https://nitter.net/Yampeleg/status/1688441683946377216
- two_in_one 3y ago> just randomly zap out some of the neurons while we train the rest It's already done: https://pytorch.org/docs/stable/generated/torch.nn.functional.dropout.html https://pytorch.org/docs/stable/generated/torch.nn.functiona...
- kaibee 3y ago> But the best cure for over-fitting is to make the dataset larger and ensure data diversity. This is also good life advice.
- BaseballPhysics 3y agoThe human brain has synaptic pruning. The exact purpose of it is theorized but not actually understood, and it's a gigantic leap to assume some sort of analogous mechanism between LLMs and the human brain.
- pcwelder 3y agoAfaik weight decay is inspired from L2 regularisation which goes back to linear regression where L2 regularisation is equivalent to having gaussian prior on the weights with zero mean. Note that L1 regularisation produces much more sparsity but it doesn't perform as well.
- nonameiguess 3y agoThis. Weight decay is just a method of dropping most weights to zero which is a standard technique used by statisticians for regularization purposes for decades. As far as I understand, it goes back at least to Tikhorov from 1970 and was mostly called ridge regression in the regression context. Normal ordinary least squares attempts to minimize the L2 norm of the squared residuals. When a system is overdetermined, adding a penalty term (usually just a scalar multiple of an identity matrix) and also minimizing the L2 norm of that biases the model to produce mostly near-zero weights. This helps with underdetermined systems and gives a better conditioned model matrix that is actually possible to solve numerically without underflow. It's kind of amazing to watch this from the sidelines, a process of engineers getting ridiculously impressive results from some combo of sheer hackery and ingenuity, great data pipelining and engineering, extremely large datasets, extremely fast hardware, and computational methods that scale very well, but at the same time, gradually relearning lessons and re-inventing techniques that were perfected by statisticians over half a century ago.
- tbalsam 3y agoL1 drops weights to zero, L2 biases towards Gaussianality. It's not always relearning lessons or people entirely blindly trying things either, many researchers use the underlying math to inform decisions for network optimization. If you're seeing that, then that's probably a side of the field where people are newer to some of the math behind it, and that will change as things get more established. The underlying mathematics behind these kinds of systems are what has motivated a lot of the improvements in hlb-CIFAR10, for example. I don't think I would have been able to get there without sitting down with the fundamentals, planning, thinking, and working a lot, and then executing. There is a good place for blind empirical research too, but it loses its utility past a certain point of overuse.
- deleted 3y ago[deleted]
- tbalsam 3y agoML researcher here wanting to offer a clarification. L1 induces sparsity. Weight decay explicitly _does not_, as it is L2. This is a common misconception. Something a lot of people don't know is that weight decay works because when applied as regularization it causes the network to approach the MDL, which reduces regret during training. Pruning in the brain is somewhat related, but because the brain uses sparsity to (fundamentally, IIRC) induce representations instead of compression, it's basically a different motif entirely. If you need a hint here on this one, think about the implicit biases of different representations and the downstream impacts that they can have on the learned (or learnable) representations of whatever system is in question. I hope this answers your question.
- joaogui1 3y agoThat looks interesting, do you know what paper talks about the connection between MDL, regret, and weight decay?
- tbalsam 3y agoI would start with Shannon's information theory and the Wikipedia page on L2/the MDL as a decent starting point. For the first, there are a few good papers that simplify the concepts even further.
- joaogui1 3y agoSorry, I know what MDL and L2 regularization are, I would like the paper that connects them in the way you mentioned
- mmmmpancakes 3y agocan you please spell out what MDL is an acronym for?
- sva_ 3y agohttps://en.wikipedia.org/wiki/Minimum_description_length https://en.wikipedia.org/wiki/Minimum_description_length