4 ms·
Show HN: Drop-in analytical replacements for standard PyTorch layers
Instead of gradient descent, each layer solves for optimal weights directly via ridge regression in a single forward pass. Works as a warm start before Adam — same final accuracy, fraction of the total training time.
Repo: https://github.com/infiplexity-pixel/to_the_point/ https://github.com/infiplexity-pixel/to_the_point/
- clbrmbr 7mo agoI’d love to see a post that clearly walks through how this works for some examples to give the intuition. And then, how much is really saved on training for a non-trivial model? And is this applicable to deep models?
- fakesum 7mo agosure: just posted on reddit: https://www.reddit.com/r/deeplearning/comments/1ro8uw2/analytical_training_for_cnns_transformers_lstms/ https://www.reddit.com/r/deeplearning/comments/1ro8uw2/analy... with benchmarks and things. and yes it works for LLMs, LSTMs, RNNs, CNNs, and more.
- clbrmbr 7mo agoThanks. Do come back to post with your tutorials. I'd recommend going quite granular and being didactic. Take somebody who understands gradient descent on MLPs and ELI5 the analytical part. Try to anticipate some of the doubts (there will be many! gradient descent from random init is dogma at this point).