4 ms·
It’s really surprising how well this works. Intuitively this illustrates how over-parameterized many LLMs are, or conversely how under-trained they might be. T
by refibrillator 3y ago
It’s really surprising how well this works. Intuitively this illustrates how over-parameterized many LLMs are, or conversely how under-trained they might be.
The drop and rescale method outlined in the paper makes the latent space increasingly sparse, which in turn allows weights to merged without much interference or degradation.
My instinct is that while merging models will have some use cases, ultimately these insights will lead to innovations in training and architecture that have the same result but with better computational efficiency.
For example instead of training an 8x7b mixture of experts then merging, just incorporate the sparsity constraint while pre-training a single 7b model (somehow).
- Drakim 3y agoDo you know if there is anybody who has made it their mission to shrink models to extremes? It feels like the sorta thing somebody would get really obsessed with doing, akin to those people who make executable out of a few bytes or shrink network payloads.
- tomxor 3y agoOoh, parameter golfing! maybe I will get into ML one day after all. On a similar note, I once found a paper that explained artificial neural networks computability from a very tiny ground up method, showing how few pieces you need to build a NAND gate... and of course once you have that you have everything.
- nico 3y agoSome time ago, Andrej Karpathy posted a YouTube video about building GPT from scratch. Using the same dataset (a 1MB text file with all of Shakespeare’s works), the video goes from a very basic model to something like GPT2 What really stood out to me is that, while the typical wisdom is “more data -> better model”, in the video they use the exact same data for all the models, so all the gains are achieved only through improved algorithms and computation power With that in mind, I wonder if at some point someone will figure out an algorithm for a model that can be trained on just an English dictionary and get a decent LLM-equivalent model that can have a basic conversation. Given the small size of the data (a dictionary), I assume the model would be pretty small and quite fast to run, even on older or smaller machines
- geon 3y agoDo you mean dictionary as just a list of words in alphabetical order, or with detailed definitions for each word?
- nico 3y agoDetailed definitions for each word, like a physical printed dictionary book
- geon 3y agoThis video? https://youtu.be/kCc8FmEb1nY?si=QR-ADr0_Y6_P2QxC https://youtu.be/kCc8FmEb1nY?si=QR-ADr0_Y6_P2QxC
- nico 3y agoYes, thank you, that’s the video
- abhgh 3y agoNot for DL models, but I was exploring doing this in a specific setting (I've posted about this earlier). For small-sized models (for some reasonable definition of size, e.g., depth of a decision tree, or # trees in a gradient boosting forest, or # non-zero coefficients in a linear model), I realized that you could make them even smaller while retaining their accuracy by selectively presenting training data to them, i.e., ignore some training data points, repeat certain others. See [1] - the x-axis shows the original model size and the y-axis shows the model size obtained by this process at the same or better accuracy. Interestingly, this meant that the conventional wisdom that the test and train distributions have to be identical for optimal held-out performance, is not true at small model sizes. It is true as models grow larger - and I was explicitly able to show this. See [2] - where the x-axis is model size, and the y-axis measures (on a scale of 0-1) how close is the optimal training distribution to the test distribution. The different lines are for different datasets. These images are from the paper here [3]. I have a library too [4]; that is in need of updates - it works today as-is though, but please use the latest minor release if this of interest [5]! [1] https://imgur.com/a/NheK49Z https://imgur.com/a/NheK49Z [2] https://imgur.com/a/N53GeNI https://imgur.com/a/N53GeNI [3] https://arxiv.org/pdf/1906.06852.pdf https://arxiv.org/pdf/1906.06852.pdf [4] https://compactem.readthedocs.io https://compactem.readthedocs.io [5] As of writing the install command should be `pip install compactem==0.9.9rc2`
- sdenton4 3y agoI tend to think there's an explore/exploit trade-off in model scale. Animals, including humans, have lots of extra neurons when they are young, which are then shed once the learning phase settles down. And this makes sense: thinking is energy intensive, so it's more efficient to sparsify. And, of course, we see a similar dynamic in ML: you can train a big model then prune it and do much better than you would by just training the smaller model directly. I've got some geometric handwaving for why this works, as well. It's easier to find a low-energy solution when you have more parameters... Sparse solutions are higher energy, and thus require longer walks (and more commitment) during training.
- sdwr 3y agoSounds about right! If you take an analogy from poker, bankroll matters - some strategies you can only play well-capitalized.
- visarga 3y ago> For example instead of training an 8x7b mixture of experts then merging, just incorporate the sparsity constraint while pre-training a single 7b model (somehow). I'm thinking it would help reduce the network demands for gradient updates if merge from time to time. That could unlock distributed training, like SETI@Home.