4 ms·
Sorry I should have been more specific - I was limiting my question to the bigger models. The smaller (~7B) models are feasible with these approaches.
by fareesh 3y ago
Sorry I should have been more specific - I was limiting my question to the bigger models. The smaller (~7B) models are feasible with these approaches.
- anon373839 3y agoAh. Perhaps the larger models will find use in in-house deployments where companies want their employees to have access to ChatGPT-like general purpose assistants, but want to prevent data from leaving their premises. LIMA shows the LLaMA 65B hitting a quality level somewhere between DaVinci-003 and GPT-4 with minimal alignment, so the base models are probably powerful enough already for this to work. Just speculating.
- MacsHeadroom 3y agoOne benefit of finetuning larger models, like 65B, is to free up limited context space vs few-shot prompting. If you want a specific kind of interaction with the model then you could take up 1/3rd of the 2048 token context window with few-shot or you could simply finetune it with QLoRA for a few hours on a consumer GPU and then get to use the full 2048 context with the finetuned model.