3 ms·
> Why would we expect to get comparable performance with models a fraction of the size and a pittance of the fine tuning? LLaMA incorporated new techniques tha
by bitL 3y ago
> Why would we expect to get comparable performance with models a fraction of the size and a pittance of the fine tuning?
LLaMA incorporated new techniques that make 65B perform way better than GPT-3's 175B so the model size argument is not very strong.