4 ms·
Yeah, the 13b model outperforms the 70b Llama 2. Goes to show how much potential there is on the software optimization front as opposed to just scaling in size
by schleck8 3y ago
Yeah, the 13b model outperforms the 70b Llama 2. Goes to show how much potential there is on the software optimization front as opposed to just scaling in size