4 ms·
Hello HackerNews! Author here :) TL;DR: We devise a linear SDE/ODE model to imitate per-class feature (thinking logits) dynamics of neural nets training based
by jiayaozhang 5y ago
Hello HackerNews! Author here :)
TL;DR: We devise a linear SDE/ODE model to imitate per-class feature (thinking logits) dynamics of neural nets training based on local elasticity (LE) [1]. We found the emergence of LE implies linear separation of features from different classes as training progresses.
The drift matrix of our model has a relatively simple structure; with that estimated, we can simulate the SDE using the forward Euler method, whose results align reasonably well with genuine dynamics.
Local elasticity models the phenomenon observed in DNN training: the effect due to training on a sample is greater for samples from the same class, and smaller for samples from different classes. For example, training an image of cats facilitates the model better learns images of other cats while not so for images of, say, dogs.
Any comments/thoughts/questions are most welcome!
[0] https://arxiv.org/abs/2110.05960 https://arxiv.org/abs/2110.05960
[1] https://arxiv.org/abs/1910.06943 https://arxiv.org/abs/1910.06943
- deleted 5y ago[deleted]
- rjeli 5y agodoes it have direct application for training nets faster? i.e. can the ode be integrated faster than backprop?
- jiayaozhang 5y ago> does it have direct application for training nets faster? That's a great question! Unfortunately not yet -- though we believe further studies may bring us there finally. We found (at least for simple classification tasks) the features seem to have a two-stage behavior: a de-randomization stage to identify the best direction in the feature space; and an amplification stage where features stretches along these directions. We've been thinking to identify a bound on the exit time of the first stage, and examine how it depends on different hyper-parameters, dataset properties etc, so that one may pinpoint how to reduce the time spending in the first stage, effectively making training faster. > i.e. can the ode be integrated faster than backprop? Also a good question, at this stage we need to estimate model parameters (the drift matrix) from simulations on DNNs. As future works we hope to explore if we can pre-determine those parameters so a comparison between backprop might make more sense.
- cinntaile 5y agoI thought the main selling point of deep learning that it finds non-linear connections in the data. Isn't it surprising that the implication is a linear separation of features?
- ShamelessC 5y agoWith regard to the class conditioned regime they are experimenting with; this is merely attempting to explain more precisely _how_ deep nets are able to distinguish features between classes. We already know that they do; but we lack a detailed model of precisely why they do and my understanding is that many base assumptions made by e.g. statistics will not help you at all with neural networks (for instance, overparameterization leading to better performance on out-of-corpus rather than overfitting).
- NumberCruncher 5y ago> many base assumptions made by e.g. statistics will not help you at all with neural networks There is lately a lot of hate against classic statistics on HN. I don't know why. Does it help to understand why and how NNs work? Not yet. But saying that it is utterly useless and won't provide any useful insights in the future sounds to me like telling the young Steve Jobs that dropping out of college and taking calligraphy classes instead of is utterly useless. And still, I am writing this on an Apple product, which set the standards for digital typography...
- ShamelessC 5y agoI'm not an expert, but statistics may be the branch of mathematics we wind up using to solve the unknowns of machine learning. I have no hate for classic statistics, sorry if I gave that impression.
- NumberCruncher 5y ago> I have no hate for classic statistics, sorry if I gave that impression. You didn't, but a lot of HNlers do. Maybe I should rant on them, that's true.
- QuantumPain 5y agoThis reminded me of Invariant Risk Minimization (IRM) (https://arxiv.org/pdf/1907.02893.pdf https://arxiv.org/pdf/1907.02893.pdf) due to the linear bound being sufficient to control the features. Do you have any comments/insight into how you’d say they’re similar/different? Thanks!
- jiayaozhang 5y agoThanks for sharing the interesting IRM paper! Will read and be back for discussion hopefully soon.
- Grieverheart 5y agoThe paper is a bit over my head. Are there any findings with respect to the phenomenon of Deep Double Descent [0], or the more recent grokking phenomenon [1]? [0] https://arxiv.org/abs/1912.02292 https://arxiv.org/abs/1912.02292 [1] https://mathai-iclr.github.io/papers/papers/MATHAI_29_paper.pdf https://mathai-iclr.github.io/papers/papers/MATHAI_29_paper....