3 ms·
How many tokens have you trained on?
by adt 3y ago
How many tokens have you trained on?
- razodactyl 3y agoAbsolutely no idea sorry!! - this was more for understanding how an LLM even works. I just happen to understand the ML pipeline so I went beyond Andrej Karpathy's tutorial. It's an extract of TinyGPT but built from the ground up to work with new features such as Multi/Grouped-Query attention, `min_p` sampling, KV-caching etc. --- Unfortunately as this is a 1-GPU personal situation everything had to be thought out from the constraints of never being able to finish training the model. --- What I can tell you though is after ~40 hours of training, the model starts showing ability to speak and perform tasks. ~80 hours it improves upon these tasks and starts showing a tiny (very minute) bit of understanding where it fits into the picture. --- GPT-2's largest was 12 layers I believe. Layers seem to correlate with ability of the model to "compute". The embedding dimension of the model seems to correlate with the ability of the model to "communicate". If you "communicate" a larger amount of information through the layers (compute) you get a much smarter AI. --- Additionally, there's an interesting effect of GQA/MQA where the model is forced to share its work between the layers. I think about it like a group of students, either working on their own or in groups: - Too many and it becomes chaos. - Too few and learning takes longer as there's no collaboration between individuals. - In my case 2 "students" per N layers means the network learns quickly and infers fast. Hopefully it makes sense, it's been a non-stop torrent of learning for myself to answer the more obvious questions that always seem to be avoided in general. Everyone is more interested (obviously) in competing and scaling up their model that the smaller models are neglected. From what I can tell, something like X.AI's Grok can be thought of as extremely high capacity without being fully saturated (it has a lot of space to learn more).
- adt 3y agoThanks. Have estimated this as 3B tokens as a round number at 8:1, but if you have a firmer number I'd love to know it. Added to the Models Table (a few 'small' and 'tiny' models there, too): https://lifearchitect.ai/models https://lifearchitect.ai/models
- razodactyl 3y agoOh wow. Thanks! Loved watching your videos on YouTube. This project had a million bug fixes and changes to the code over time so there have been hundreds of models that have been trained / deleted. My methods aren't exactly scientific unfortunately. I'll do a write up one day but for anyone else interested I have uploaded some information here https://ftp.bytebreeze.dev/ftpuser/ https://ftp.bytebreeze.dev/ftpuser/ I need to do a serious cleanup to the code and host it somewhere more stable but the raw version of what I have so far can be found there. If anyone gets it running, keen to hear feedback!