5 ms·
Do you know if there is anybody who has made it their mission to shrink models to extremes? It feels like the sorta thing somebody would get really obsessed wit
by Drakim 3y ago
Do you know if there is anybody who has made it their mission to shrink models to extremes? It feels like the sorta thing somebody would get really obsessed with doing, akin to those people who make executable out of a few bytes or shrink network payloads.
- tomxor 3y agoOoh, parameter golfing! maybe I will get into ML one day after all. On a similar note, I once found a paper that explained artificial neural networks computability from a very tiny ground up method, showing how few pieces you need to build a NAND gate... and of course once you have that you have everything.
- nico 3y agoSome time ago, Andrej Karpathy posted a YouTube video about building GPT from scratch. Using the same dataset (a 1MB text file with all of Shakespeare’s works), the video goes from a very basic model to something like GPT2 What really stood out to me is that, while the typical wisdom is “more data -> better model”, in the video they use the exact same data for all the models, so all the gains are achieved only through improved algorithms and computation power With that in mind, I wonder if at some point someone will figure out an algorithm for a model that can be trained on just an English dictionary and get a decent LLM-equivalent model that can have a basic conversation. Given the small size of the data (a dictionary), I assume the model would be pretty small and quite fast to run, even on older or smaller machines
- geon 3y agoDo you mean dictionary as just a list of words in alphabetical order, or with detailed definitions for each word?
- nico 3y agoDetailed definitions for each word, like a physical printed dictionary book
- geon 3y agoThis video? https://youtu.be/kCc8FmEb1nY?si=QR-ADr0_Y6_P2QxC https://youtu.be/kCc8FmEb1nY?si=QR-ADr0_Y6_P2QxC
- nico 3y agoYes, thank you, that’s the video
- abhgh 3y agoNot for DL models, but I was exploring doing this in a specific setting (I've posted about this earlier). For small-sized models (for some reasonable definition of size, e.g., depth of a decision tree, or # trees in a gradient boosting forest, or # non-zero coefficients in a linear model), I realized that you could make them even smaller while retaining their accuracy by selectively presenting training data to them, i.e., ignore some training data points, repeat certain others. See [1] - the x-axis shows the original model size and the y-axis shows the model size obtained by this process at the same or better accuracy. Interestingly, this meant that the conventional wisdom that the test and train distributions have to be identical for optimal held-out performance, is not true at small model sizes. It is true as models grow larger - and I was explicitly able to show this. See [2] - where the x-axis is model size, and the y-axis measures (on a scale of 0-1) how close is the optimal training distribution to the test distribution. The different lines are for different datasets. These images are from the paper here [3]. I have a library too [4]; that is in need of updates - it works today as-is though, but please use the latest minor release if this of interest [5]! [1] https://imgur.com/a/NheK49Z https://imgur.com/a/NheK49Z [2] https://imgur.com/a/N53GeNI https://imgur.com/a/N53GeNI [3] https://arxiv.org/pdf/1906.06852.pdf https://arxiv.org/pdf/1906.06852.pdf [4] https://compactem.readthedocs.io https://compactem.readthedocs.io [5] As of writing the install command should be `pip install compactem==0.9.9rc2`