3 ms·
I did use a tweaked nanoGPT to pretrain a 12M model on TinyStories (2Gbytes produced by GPT4), and results are pretty amazing. I've adapted it a bit on Wikipedi
by novaRom 3y ago
I did use a tweaked nanoGPT to pretrain a 12M model on TinyStories (2Gbytes produced by GPT4), and results are pretty amazing. I've adapted it a bit on Wikipedia then, and it looks like a solid bullshit generator, much smarter than any smoothed n-gram model, and significantly smaller. My bet small LLMs will be predominant in multiple areas. My next goal is to reduce 7B llama2 to 10-100M without making it much dumber.
- GaggiX 3y ago>My next goal is to reduce 7B llama2 to 10-100M without making it much dumber. That is going to be hard as the 7B model was trained on 2T tokens. Maybe if you heavily restrict the range in which the model should operate.
- ljlolel 3y ago1. It’s faster and cheaper to train a smaller model 2. Better than tokens is to train on probability distributions (distillation) and trees of probability distributions
- nickpsecurity 3y agoI've never seen anything about training on probability distributions or trees of them. Do you have articles with examples you could share with us? I did try a quick search for it. Found some interesting papers. The links to them are below in case anyone finds them interesting. https://arxiv.org/abs/2212.11481 https://arxiv.org/abs/2212.11481 https://towardsdatascience.com/a-new-way-to-predict-probability-distributions-e7258349f464 https://towardsdatascience.com/a-new-way-to-predict-probabil... https://arxiv.org/pdf/1912.07913.pdf https://arxiv.org/pdf/1912.07913.pdf https://dukespace.lib.duke.edu/dspace/bitstream/handle/10161/25822/Awaya_duke_0066D_16951.pdf?sequence=1 https://dukespace.lib.duke.edu/dspace/bitstream/handle/10161...
- Remmy 3y agoWould love to read more about your time in NanoGPT. I've been getting familiar with it myself lately and it's still pretty much gibberish in the output with 16M, but the dataset is admittedly trash right now as well.
- oaguy1 3y agoI also trained nanoGPT on TinyStories, produced about a 32M model. The results are amazing, especially considering I opted for a character-level model similar to the toy dataset in the repo. I’m writing about the experience while also doing a deep dive into the code on medium (username oaguy1). Smaller LLMs are definitely worth considering with the right quality training data. Once I finish playing with TinyStories, I recently tweaked the Standardized Project Gutenberg Corpus (~11GB) to be more modern. Want to see what I can do with it with nanoGPT and then maybe Huggingface’s libraries.
- hekec 3y agoHow do you adapt it on Wikipedia? Do you just add it to the dataset and continue training?