3 ms·
One finding in the LLaMA paper [1] is that our current large models are undertrained. LLaMA with 13B params outperforms GPT-3 175B (not ChatGPT), but an "instru
by v64 4y ago
One finding in the LLaMA paper [1] is that our current large models are undertrained. LLaMA with 13B params outperforms GPT-3 175B (not ChatGPT), but an "instruct" version of LLaMA was finetuned over the 65B model and did quite well.
[1] https://arxiv.org/pdf/2302.13971.pdf https://arxiv.org/pdf/2302.13971.pdf