4 ms·
Hah funny to see this on HN, it is a relatively old project but one that I continue to love and still work on. I was trying to train a GPT one day and discovere
by karpathy 4y ago
Hah funny to see this on HN, it is a relatively old project but one that I continue to love and still work on. I was trying to train a GPT one day and discovered that available implementations were quite complex, spread across many files, and took way too many kwargs switches for esoteric/rare options that just bloated and complexified the code. But in my head a GPT was a super simple neat, isotropic model, so I got all worked up and wrote minGPT.
The project went on to have more impact than I originally imagined and made its way into a number of projects and papers. One of those I found only a few days ago here: https://twitter.com/karpathy/status/1566100736076697600 https://twitter.com/karpathy/status/1566100736076697600 . What I love about these projects is that the authors often "hack up" minGPT in code directly. They don't configure a comprehensive kwarg monster. I think there's a beauty in that. Very often I wish we had more gists and fewer frameworks - to look at code chunks, understand them completely, tune them to our needs, and re-use them in projects, similar to how bacteria trade little DNA plasmids. minGPT is written for those who want that for their GPT projects. There's plenty of cons to this approach too, ultimately I think there's value in both approaches.
Coming up the theme of future minGPT development: more examples, and more teeth - it should be possible to demonstrate the training of relatively serious (~few B) models with minGPT on one n-gpu node and reproduce some benchmarks around that scale, but never sacrifice its readability.
- mmulet 4y agoThanks for making it it! There is immense value in something you can just dive into and hack on. I’ve been hacking on stable Diffusion/latent diffusion these past couple weeks, and you don’t know how much time it would have saved me, if it just had something similar!
- jphoward 4y agoI completely agree! I personally find these powerful new network releases border on the depressing, in that they aren’t really network releases but huge training systems of dispersed YAMLs. YOLOv4 was a case in point where I was too overwhelmed to try and integrate it into a project I was working on. PS you are a hero of mine - I’m an academic medical doctor for who CS231n was my first foray into AI, and since then I’ve gone on to gold medal in a couple of Kaggle competitions and secured 5 years of higher research funding to pursue clinical AI. I am immensely grateful to you and Fei-Fei Li.
- darawk 4y agoFor anyone else who was new to the phrase "isotropic model": https://github.com/christianversloot/machine-learning-articles/blob/main/introduction-to-isotropic-architectures-in-computer-vision.md https://github.com/christianversloot/machine-learning-articl...
- albertzeyer 4y agoThis works for an architecture which has been well tuned and studied before, like LSTM or Transformer. Once you do research on the model, testing out things, it often tends to become such kwarg monster in many frameworks. Having everything (relevant) in one file (even in the config file itself with hyper params) allows you to copy the file for every experiment and modify it inplace. This avoids the kwargs mess. But then the config files are very complex, and can become messy in other ways (esp for research projects). Example: https://github.com/rwth-i6/returnn-experiments/blob/master/2020-rnn-transducer/configs/rna3c-lm4a.convtrain.switchout6.l2a_1e_4.nohdf.encbottle256.attwb5_am.dec1la-n128.decdrop03.decwdrop03.pretrain_less2_rep6.mlr50.emit2.fl2.fixmask.rna-align-blank0-scratch-swap.encctc.devtrain.retrain1.config https://github.com/rwth-i6/returnn-experiments/blob/master/2... Such approach makes it much more flexible and does not mess with the baseline code. As you say, it's more like an evolutionary DNA-like approach, where you then tend to do crossovers with other evolved good-performing configs, etc.
- porphyra 4y agoIn Python I have found that a good way to deal with large config files is to use Dataclasses and serialize or deserialize them with OmegaConf.
- Siira 4y agoAre there any similarly structured projects around?
- frozencell 4y agoWhat's required for an AGI to update karpathy's HN bio?