4 ms·
That might be true for finetuning ChatGPT 3.5, but if you can finetune a small model (7B or less) to perform on par with GPT-4, while being faster and private,
by thorum 3y ago
That might be true for finetuning ChatGPT 3.5, but if you can finetune a small model (7B or less) to perform on par with GPT-4, while being faster and private, that’s a different story.
- smallnamespace 3y agoYou definitely can't in the general case (for example, your 7B model is never going to be able to help much with coding, fine tuning or no). It can make sense if you have a particularly simple use case.
- qeternity 3y agoBy definition you wouldn’t fine tune a 7B model to be generally as good at GPT4. You would just be trying to overfit some small amount of functionality in a narrow domain.
- smallnamespace 3y agoYes but from the context of this discussion, we’re trying to figure out the “sweet spot” model size where it’s worth attempting fine tuning. My guess is it’s only worthwhile for matching simple tasks with small models, and any sufficiently complicated task it’s better to do few/zero shot instead.