4 ms·
I’d like some clarification of terms - when they say it takes 3 hours to train, they’re not saying from scratch are they? There’s already a huge amount of train
by _9hey 4y ago
I’d like some clarification of terms - when they say it takes 3 hours to train, they’re not saying from scratch are they? There’s already a huge amount of training to get to that point, isn’t that correct? If so, then it’s pretty audacious to claim they’ve democratized an LLM because the original training likely cost an epic amount of money. Then who knows how much guidance their training has incorporated, and it could have a strong undesirable viewpoint bias based on the original training.
- joshhart 4y agoThe 3 hours is the instruction fine-tuning. The base foundational model is GPT-J which was already provided by Eleuther-AI and has been around for a couple of years. Note: I work at Databricks and am familiar with this project but didn't work on it.
- Taek 4y agoDo you know why GPT-J is being used instead of NeoX or any of the other larger open source models?
- joshhart 4y ago7B is a sweet spot where you can do something with limited resources both for training and inference. Going beyond that you spill out of an A100 without tricks. We will continue iterating on this with other models.
- dragonwriter 4y agoIf fine tuning a small model, which can be run on consumer hardware once trained, provides quality results, why use a larger base model?