4 ms·
This is a lot of tokens. Llama 2 was trained on two trillion tokens [1] [1] https://arxiv.org/abs/2307.09288 https://arxiv.org/abs/2307.09288
by tydunn 3y ago
This is a lot of tokens. Llama 2 was trained on two trillion tokens [1]
[1] https://arxiv.org/abs/2307.09288 https://arxiv.org/abs/2307.09288
- amilios 3y agoLoss was still decreasing for the models, there's a sense that we can push the training data much much further than we currently are.
- npsomaratna 3y agoYup. I found this article quite enlightening: https://espadrine.github.io/blog/posts/chinchilla-s-death.html https://espadrine.github.io/blog/posts/chinchilla-s-death.ht...
- rushingcreek 3y agoPhenomenal blog post about scaling laws.
- deleted 3y ago[deleted]
- famouswaffles 3y agoPrediction as an objective basically forces the models to model the casual processes that create the text itself. It's not going to stop getting better unless the data is insufficient/unvaried or the architecture creates a bottleneck. I think by the time the former is an "issue", we'll have a Super Intelligence on our hands anyway. The latter is looking less and less likely to be a real hurdle. Very little inductive bias to steer away from crucial solutions, very scalable.
- yorwba 3y agoThe TinyLlama project is trying to do that pushing by training a small 1.1 billion-parameter model on 3 trillion tokens: https://github.com/jzhang38/TinyLlama https://github.com/jzhang38/TinyLlama
- minimaxir 3y agoAnd Llama 2's training data was less aggressively deduplicated.