5 ms·
The headline and introduction on the linked page say "You can now train a 70b language model at home. We’re releasing an open source system, based on FSDP and Q
by qsi 3y ago
The headline and introduction on the linked page say "You can now train a 70b language model at home. We’re releasing an open source system, based on FSDP and QLoRA, that can train a 70b model on two 24GB GPUs."
How does "fine tuning" differ from "training?" Reading the linked article I had assumed I could create my own trained LLM at home with two 24GB GPUs.
- jph00 3y agoThe article actually sneaks in a footnote that answers this (https://www.answer.ai/posts/2024-03-06-fsdp-qlora.html#fn1 https://www.answer.ai/posts/2024-03-06-fsdp-qlora.html#fn1): "Throughout this article “training” can refer to either pre-training, or fine-tuning". (Generally, we've told students at fast.ai since 2017 that they should almost never be starting from random weights -- most of the time it's best to start with a pretrained model and fine-tune that, even if it's from a somewhat different domain to the problem you're working on.)
- Tomte 3y agoHave you changed your mind on „The End of Finetuning“ (https://www.latent.space/p/fastai https://www.latent.space/p/fastai ) or did I simply misunderstand that? Oh, and thanks for quirky stuff like your APL video!
- jph00 3y agoThe title of that podcast isn't something I actually said (IIRC). I commented in that interview that I feel we should not consider pre-training and fine-tuning to be as separate as we do now.
- Tomte 3y agoSo you‘re generally in favor of mixing training data without separating them in phases, but when I use pretrained weights (as you recommend instead of random weights) I generally do not have access to whatever the neural net was pretrained with by someone else, so I have to make do with my finetuning data, yes? Thank you!
- pama 3y agoYes.
- swyx 3y ago"The right way to fine-tune language models... is to actually throw away the idea of fine-tuning. There's no such thing. There's only continued pre-training." :) i hope i didnt pervert your intent too too much for clickbait or something, i thought it was the spirit of what you said
- keremturgutlu 3y agoYou most definitely can, the main difference is that only partial ~2% of the parameters get updated during training. Say you start from a model like llama-70B which already knows english and has some world knowledge based on its pretraining dataset. It might not be ideal for drastic domain shifts, such as adapting a model to learn new languages (which might require a new tokenizer and model embeddings) but still might be possible to some extent.
- qsi 3y agoThank you for clarifying. I have been wanting to dip my toes into LLMs at home but obviously I have a steep learning curve ahead of me, and would need considerably beefier hardware!
- chasd00 3y agoIt’s steep but manageable, absolutely go for it. The more people who understand the tech the better.
- IanCal 3y agoYou can take an existing 70B model and train it to do a more specific task. You're teaching it the task but you're relying on a foundation model for the base understanding of the world/words/etc.
- qsi 3y agoOK, that makes sense. Thank you!