3 ms·
I felt like the point of llama 3 was to prove this out. 2T -> 15T tokens, 70B -> 405B. We now have llama 3.3 70B, which by most metrics outperforms the 405B mo
by deepsquirrelnet 2y ago
I felt like the point of llama 3 was to prove this out. 2T -> 15T tokens, 70B -> 405B.
We now have llama 3.3 70B, which by most metrics outperforms the 405B model without further scaling, so it’s been my assumption that scaling is dead. Other innovations in training are taking the lead. Higher volumes of low quality data aren’t moving the needle.