4 ms·
it lags behind because according to their blogpost it was trained on <300B tokens. LLaMAs as far as I know were trained on more than trillion
by option 4y ago
it lags behind because according to their blogpost it was trained on <300B tokens. LLaMAs as far as I know were trained on more than trillion
- gpm 4y agoThe LLaMa paper says 1 trillion for the smaller models (7B, 13B) and 1.4 trillion for the larger models (30B, 65B)
- deleted 4y ago[deleted]