7 ms·
A Theory on Adam Instability in Large-Scale Machine Learning
- bacf 3y ago[flagged]
- claytonjy 3y ago[flagged]
- deleted 3y ago[deleted]
- optimalsolver 3y agoIs it possible to use derivative-free/black-box optimizers to train these large networks? From what I understand, gradient descent and its cousins can't suddenly jump to a distant global optimum.
- tomrod 3y agoparticle swarm optimization, genetic algorithms, and tabu search/heuristic search are some items I'm aware of to force out of local optimum. Using Halton sequences can also help cover the space for search initialization, versus simple random draws in a space.
- zerodensity 3y agoFrom my granted limited understanding Adam is basically gradient decent combined with heuristic search.
- yobbo 3y agoAdam is somewhat analogous to an audio compressor on the gradient "signals". (Edit: eh sort of ...)
- perone 3y agoNo.
- dontwearitout 3y agoMore correctly: we don't know how to do this efficiently. Biological neural networks don't use backpropagation and work great.
- sdenton4 3y agoConference takes years, though... It's entirely possible that back prop is unrealistic but far more efficient than biological learning.
- ShamelessC 3y agoOut of my depth so happy to be corrected. Don’t many/most state of the art models take many months to train on far more data than humans need for similar tasks? Also, while e.g. GPT4 is quite capable across many tasks - humans seem to average towards learning robust _learning techniques_ themselves. Learning a new subject becomes easier thanks to somehow tracking and encoding learning strategies that are robust to learning other unrelated topics.
- landryraccoon 3y ago> Don’t many/most state of the art models take many months to train on far more data than humans need for similar tasks? Humans generally need 18 years of pre training followed by 4-6 years of fine tuning before they can “one-shot” many difficult tasks. That’s way more training than any machine learning model I’m aware of. Even for tasks like reading the newspaper and summarizing what you read, you probably had to train for 10-12 years.
- tacheiordache 3y agoI see this stance of yours parroted over and over but a 3 year old can tell a dog from a cat doesn't need to be trained on millions of images. Also uses way less energy for that.
- NavinF 3y agoPeople have tried, but all the local minimums perform about the same so there's no point trying to find a global minimum. A much better strategy is to train multiple models and use all of them at inference time to score better
- PartiallyTyped 3y agoFirst order optimizers have trouble because they fall into local minima, however, in practice things are different. When your parameter space is in the order of billions, for all practical purposes, there is always a direction of descent. More over, local minima seem to be rather close to the global minima.
- peheje 3y agoI believe there are still ongoing efforts in this. See for example the forward forward algorithm. https://www.google.com/url?sa=t&source=web&rct=j&opi=89978449&url=https://www.cs.toronto.edu/~hinton/FFA13.pdf&ved=2ahUKEwjYu5b31ZiAAxXrRPEDHVMoC1QQFnoECA8QAQ&usg=AOvVaw1_CmnqC9TpQjiXftmaP2sz https://www.google.com/url?sa=t&source=web&rct=j&opi=8997844...
- esafak 3y agoYou have a gradient so use it instead of faffing about. As another user said, the optima are all the same since the models are wildly over-parameterized.
- mturmon 3y agoThis is tersely stated, but it's wise. In general, following the gradient (if you have it) is a very, very good idea.
- Der_Einzige 3y agoYes, and the only reason it doesn't work is that no one has written truly fast, GPU implementations of them. Don't let anyone here teach you otherwise, even small scale crappy versions (like what I could code in numpy) can successfully solve reinforcement learning problems rather quickly. Nay-sayers might tell you that it doesn't work, but they are wrong. Global optimization is strictly superior to local optimization in general, and we in the AI field are stuck deep in a local minimum right now. Here's me implementing an algorithm from 2009 in single-core on a CPU and getting pretty excellent results on RLHF benchmarks: https://github.com/Hellisotherpeople/Python-Cooperative-Synapse-NeuroEvolution https://github.com/Hellisotherpeople/Python-Cooperative-Syna...
- marcinzm 3y agoThe whole argument is that at large scales with billions of params it doesn’t matter specifically because of those billions so giving a toy example seems to miss the point.
- adeon 3y agoI have actually attempted this recently. I took a small 10M parameter Shakespeare language model used as an example in nanoGPT, swapped out gradient descent, tested various black-box optimizers from what I could find in literature. It takes 3 minutes to train the Shakespeare model with gradient descent. The black-box methods I tested so far likely take 30+ hours to train (I haven't tried to take them to the end yet). I've hit a wall where progress is very slow. The text generated at that stage has punctuation and words are split with spaces but the words themselves are mostly nonsense. Almost feels like it learned that English is letters separated by spaces, and that you put exclamation marks or periods at the end but not that much more. There's some larger scale CMA-ES variants I still want to test that don't have quality implementations. I've tried to stare at pictures of gradients and weights from half-trained models and trying to come up with ideas how to get there with black-box optimization. Also trying some original ideas where you compute a gradient, but you would not compute it against a loss function. The gradient would be more for discovering hidden structure in weights, that you would then put on some black-box optimizer as a guide (which I guess makes it not entirely black box. Gray box?) Possible? I mean, I guess technically. Practical? No way, unless some major breakthrough happens. My current goal is to just produce a model, even if training takes laughably long so I can say I've trained a language model using nothing but getting a fitness score from a black box function. Edit: if you are reading this and are aware of any other serious attempts at training a non-trivial sized language model without gradient descent I would want to know. So I can try their methods. I know there's some large scale stuff used in reinforcement learning like in one Uber paper but not in LLMs specifically.
- dadoomer 3y agoVery interesting, thanks for sharing. I would be interested in reading more on gradient-free optimization applied to large problems (like LLMs).
- Legend2440 3y agoNo one has the scale to make that happen. It's about information. Gradient-free methods integrate little or no information about the problem; they're a blind watchmaker. This works, but it's slow and gets slower the bigger your problem is. (curse of dimensionality) Gradients integrate some limited information about the problem. This lets you find solutions much faster, and neural networks are structured specifically to be easy to optimize with gradients. Local minima don't seem to be a problem. The future is probably even smarter optimizers that integrate more information about the problem and learn to make good assumptions. This is the goal of Learned Optimizers, like Velo (https://arxiv.org/abs/2211.09760 https://arxiv.org/abs/2211.09760).
- itsthecourier 3y agoSo it happens that when training LLM some training batches worsen them model, but such batches actually improve it when fed later, why?
- itsthecourier 3y ago"In this work, we argue that the training loss instabilities observed in large-scale training should be associated with the time-domain correlation between the gradient estimates of earlier layers in the deep-learning models. Based on the identified connection, we propose several ways to mitigate the instabilities, along with the heuristic method that was known in the literature. We conclude that at this point, there is no silver bullet to solve the problem, and the appropriate remedy depends on the specific setup of the large-scale training run."
- js8 3y agoSo, it's a form of superstition?
- mannykannot 3y agoThis may be naive, but the gradient seen during one training batch would not depend only on the content of that batch, but also the outcome of all previous batches (or so I suppose.) If that is so, then whether one of these spikes occur is not only a function of the batch, but also the sequence of prior batches.
- typon 3y agoIt makes intuitive sense. If I speak Japanese at you (and you are a non-Japanese speaker), I will just confuse you. Instead if you spend two years learning Japanese, then I share some information with you in Japanese, you will learn something new and become more knowledgeable.
- K0balt 3y agoThat is a good analogy. The insight is improved by realising that in the human context the confusion is temporary and results in the rejection of the data. In the LLM it is forced into the matrix in the incorrect context, so it is harmful.
- yobbo 3y agoWhen gradients become auto-correlated, they are not zero mean. In that, case vₜ can grow large and uₜ becomes very small. Maybe try centering the g² term in vₜ like (gₜ-mₜ)². Also, when restarting training, m and v are probably reset to zero. Might not be related to the phenomena in the paper.
- troelsSteegin 3y ago"The name Adam is derived from adaptive moment estimation.", https://arxiv.org/abs/1412.6980 https://arxiv.org/abs/1412.6980
- hoseja 3y agoThis is not a coincidence, because nothing is ever a coincidence.
- deleted 3y ago[deleted]