3 ms·
They list 7B, 13B, 33B, 65B architectures. Presumably, they compare 65b one to GPT-3 175B. Chinchilla model which is about 70B outperformed a much larger GPT-3
by rnosov 4y ago
They list 7B, 13B, 33B, 65B architectures. Presumably, they compare 65b one to GPT-3 175B. Chinchilla model which is about 70B outperformed a much larger GPT-3 model. So not that fantastical.
EDIT: I stand corrected. They do compare 13B model with a large GPT-3 model which is hard to believe without a bit more concrete evidence
- rileyphone 4y agoYou're incorrect. > For example, LLaMA's 13B architecture outperforms GPT-3 despite being 10 times smaller. Of course this is all on benchmarks but it's a big improvement if true.