6 ms·
Llama 3 8B basically answers the question of: "What happens if you train a small(-ish) model for a very long time". I think we've seen this trend start with so
by fnands 2y ago
Llama 3 8B basically answers the question of: "What happens if you train a small(-ish) model for a very long time".
I think we've seen this trend start with some of the Mistral models going beyond the Chinchilla optimal point, and now with Llama 3 even further. 15T tokens for a 8B param model is a lot more than we've seen so far (for context, Llama 2 was 2T tokens), and it seems to be paying off.
If anything, this release makes me excited for the quality of smaller models going forward.
- Tenoke 2y ago>If anything, this release makes me excited for the quality of smaller models going forward. While there's going to be more that can be milked out of them, we are already clearly fairly deep in the diminishing returns portion on small models.
- MyFirstSass 2y agoSeems like it when for 7b models: (but also wow things are moving fast!); Llama - 1.4 Trillion tokens - feb. 2023, Llama 2 - 2 Trillion tokens - july. 2023, Mistral - 8 Trillion tokens - sep. 2023 (this was the first big impressive leap where local models really became useful for chat) Llama 3 - 15 Trillion tokens So we went 1.4x, to 4x, to 2x. I wonder if there's even unused data out there? Haven't tried the new model enough to see how much better it is than Mistral, that will be the real SOTA test for now.
- Escapado 2y agoI wondered about the same thing but at the same time only about 5% of the training data is non-English and I would be surprised if the total amount of published text from all non English languages combined was also just 5% of all published text. So my intuition tells me there is still heaps of data but what might be tricky is to properly access and asses it’s quality. Also the redpyjama v2 dataset has 30T tokens and is based on common crawl. Now I don’t know too much about common crawl but I doubt it has in it all published scientific books in all the different languages as these are often not freely crawlable. I remember when I studied physics there were at least 20 different 200-800 page long books on particle physics in German alone in our campus library. That must amount to 5million token by itself, from just one niche of physics. The Hamburg public library hosts about 5 million books and 90 million scientific articles mostly in English and German. If the average length of a scientific article is 6000 tokens and the average book about 100000 then that alone is already 1 trillion token. I bet there are significantly larger libraries and this is before even crawling the internet and looking at other languages or even generating training data.
- sheepscreek 2y agoStill quite impressive to think these models are trained on 15x the content in a large public library. That is insane.
- vidarh 2y agoDeutsche Nationalbibliothek appears to have 43.2 million "items", of which apparently 17.3 million are books. If we assume ~60,000 tokens for an average book, which seems very conservative given average word length in German and a novel typically being considered anything above ~40k words), that's another trillion just for their books, so I'm guessing the total German language content available in major libraries will be many times that. E.g. the Norwegian National Library has somewhere between 3x and 10x as many tokens in Norwegian newspapers as in books (at one point I think GPT3 breakdown of training data by language surfaced, and the Norwegian data was a tiny fraction of what was available in the national library, even before trying to estimate online/digital content). While I'm sure there's overlap [1] between the languages, a lot of it will help translation, and I think even for smaller languages the ratio of local content seems to dwarf translations. E.g. the "bestsellers" from English, French, and German all get translated to Norwegian, but most of the "long tail" content is local. [1] I was tickled to a find one of my uncles represented in Deutsche Nationalbibliothek; he was a professor in statistics, so it was a translation of some of his research
- londons_explore 2y ago> I wonder if there's even unused data out there? One day someone is going to train on the contents of DM's/private conversations/emails. There has to be 50x or more the quantity of that compared to public text. I suspect they'll do it via some 'prove-ably private training' regime, and therefore be able to claim it isn't a privacy violation.
- nolok 2y agoYes, they're either already getting into it behind legal facade, or aiming for it as the next eldorado of data. Facebook has whatsapp and messenger, microsoft has skype and outlook and exchange and msn messenger, google has gmail and all your text message and their bazillion chat apps and usenet and irc and ..., apple has imessage and icloud emails and ... There is so much data there, it dwarfs those token count.
- londons_explore 2y agoThe players with e2e encryption (imessage, whatsapp) would need to do some kind of client side edge device training. Possible, but hard to do with nobody knowing, and edge device training usually involves big quality compromises.
- nolok 2y agoDidn't both of them have some of their backup in clear text ? I know whatsapp backup on gmail were.
- Workaccount2 2y agoGoogle is sitting on ~20 years of gmail, but I can imagine the headache of both cleaning the dataset and likely consumer blow back. They also have youtube, which almost certainly has enough good data to train a powerful model on it's own, but also seems daunting to clean up first.
- seunosewa 2y ago
- nolok 2y ago> I wonder if there's even unused data out there? I think you're massively underestimating the amount of data out there. The challenge is how to access and categorize that data. Every usenet message, forum post from old school bbs to php forums to modern javascript abomination and closed discord boards, every email, every text message, every irc message, ... Those are probably a pain point to access due to rules and regulation and yada yada, but that alone dwarfs the 15 trillions, and you've not even started on actual quality content. (not saying these would specifically be good for llm, just answering to your "unused data" assessment)
- MyFirstSass 2y agoThat's a good point, though i already thought the OpenAI team had been very agressive in sweeping both reddit, usenet, + various illegal megatorrents of books, forum dumps etc. I remember there were some controversy around it on twitter a few months ago. One thing though is books/content/media from other language spheres though that could probably at least 10x the size of the data, and as far as i know translation starts to work rather well in these larger models so it would probably just plug right into the knowledgegraph for all languages?
- vidarh 2y agoThere are still vast amounts of data locked up behind login screens etc., though. E.g. to the foreign language data, a lot of national libraries around the world are either not even fully digitized yet or have lots of locked-down content. The Norwegian one is pretty open, but there's still huge amounts (like most newspapers newer than a century or so) that is either only available based on geolocation (I have my VPN for genealogy because of that - I'm Norwegian but live in the UK, and it's a nuisance), or only in a physical library in Norway. Similarly I was looking for something from the British Library at one point and it was behind a paywall (a copying fee). I have no idea how to even start to estimate how much data is locked down like that, and it's harder yet to try to figure out which parts of that it'd be possible to negotiate access to for various players, and what they can circumvent (e.g. say by buying book collections and the like - OpenAI is large enough by market cap it could afford to buy some of the largest extant publishers, for example, if they thought it gave them sufficient benefits).
- CuriouslyC 2y agoAre we though? We haven't even started training 1.58b models, synesthetic data and "model gyms" are very promising and new architectures have been coming out that offer real benefits over transformers.
- kleiba 2y agoIsn't overfitting the usual result of training a model for a long time?
- bjornsing 2y agoAs I understand it the Chinchilla “optimal point” is severely under-fitted. It’s optimal in the sense that if you only care about training cost it would have been better to make the model bigger and even more under-fitted. But clearly we care about inference cost too (or even primarily), so it makes sense to train for longer. Also, these models are trained on trillions of tokens, so I’m not sure an 8B model even can overfit.
- candiodari 2y ago1) obviously, when the people doing the inferencing are different from the ones doing training, training cost does not matter to inference. 2) Almost, Chinchilla concerns itself with minimizing cost(training) + cost(inference). The expectation most people have is that there's an insane amount of inference compute, and training is maybe 1% of that. The point of the Chinchilla paper is that that's not true, training uses such insane amounts of compute that despite all model inference by half the internet for a year or two is a huge amount of compute, training is still a very decent percentage of that. I believe in one of the examples they pointed out that even the whole internet inferring with a model for years was still only 20% of the cost of training that model. People expect it works like compilers, that making a compiler produce 1% faster code is worth 50 highly-paid SWEs because while an individual program run isn't exactly expensive, the time and resources spent running programs is astronomically larger than the time and resources spent developing compilers. The thing is most of the optimizations we know don't work during training. You can't quantize, you can't MoE (well, you can, obviously, but it doesn't save any training computation. In fact it increases training cost) At 50-50, having a 10% cheaper-to-train model justifies 10% more expensive inference. 3) Combining both arguments ... at this point people should probably realize that Facebook's LLama is really an attack on Google (which is at least partially working, elon musk is tweeting about it) If Facebook really doesn't care about AI (or ... about as much as, say, netflix does. Not zero, but as long as they beat reddit's efforts they feel very comfortable), but Zuck does care about destroying Google, the calculus changes. Zuckerberg may not want the best possible AI, he may want as many scammers as possible trying to Game the Google search quality team, to present them with challenges faster than they can adapt. Then the training cost becomes a moot point. Hmmm, I should send my CV to meta ...
- machiaweliczny 2y agoThere's still lot's of low hanging fruit in data preparation for these models it seems. I wonder if we will reach insane quality when we will be able to allow models to generate not exact same text but something similar (thus allowing for more compression). I think math leans a lot to this type of training as one can super easily verify if generated result is fine or not. We probably should do similar thing for language when other model judges if generation is ok or not (but doesn't require exact generation). This probably can be done as fine tuning? Anyone know if this has been tested already?
- threatripper 2y agoCould this be implemented by making the loss zero if the output for the expected token is above some threshold? E.g. if the threshold is 0.1 it can output 10 likely tokens with 0.1 probability each and if any of them is correct it will not modify the weights.
- wongarsu 2y agoOne issue is that this doesn't punish the model of the top prediction produces complete nonsense. I can imagine getting good results when fine tuning an existing model though
- wongarsu 2y agoSo something like having a smaller model generate embeddings, and use the embedding distance between the predicted text and the expected text for the loss computation? That sounds like it would be incredibly expensive. But maybe it allows you to fine tune on less data, offsetting the cost a bit (and allowing you to fine tune on just your very best data)
- zitterbewegung 2y agoYou could derive similar data by looking at the embedding and then synthesize similar data and then add it to the training dataset. Right now fine tuning a model before synthesizing the similar data could be done.
- 2y ago
- deleted 2y ago[deleted]
- freehorse 2y agoI wonder if quantisation in these models actually reduces the quality more than the less overtrained ones we were used to? Anybody knows if that could be the case?
- WithinReason 2y agoThat would make sense
- imjonse 2y agoToo bad they did not train or release a Llama 3 2B to see how it fares against Phi and Gemma.
- cchance 2y agoim surprised heirs no 14b or 16b
- Aissen 2y ago> very long time On one of Meta's 24k H100 clusters running at 95% efficiency, it's 2.3 days.
- WithinReason 2y agoIt was closer to 40% efficiency AFAIK
- Aissen 2y agoI'm curious, do you have any pointers? This is what this article mentions: > Those improvements resulted in an overall effective training time of more than 95% https://ai.meta.com/blog/meta-llama-3/ https://ai.meta.com/blog/meta-llama-3/
- logicchains 2y agoThe known SOTA for GPU flops utilisation for training on that many GPUs is somewhere between 50-60%, e.g. https://github.com/NVIDIA/Megatron-LM https://github.com/NVIDIA/Megatron-LM . If they really managed to get 95%, it's a huge advance in the state of the art. I guess by training time they meant with respect to downtime of the hardware, not utilisation of the hardware potential.
- Aissen 2y ago> I guess by training time they meant with respect to downtime of the hardware, not utilisation of the hardware potential. That was my understanding as well, those are two different levels of efficiency.
- WithinReason 2y agoThat means that 5% of the time GPUs were not doing any training, e.g. the server was busy with checkpointing so the training paused. 40% is according to Karpathy: https://twitter.com/karpathy/status/1781028605709234613 https://twitter.com/karpathy/status/1781028605709234613
- jasonjmcghee 2y agoZuck mentioned they also cut entire classes of data from training Llama 2 that they included in Llama 3 as Llama 2 was intended to be used for social / meta related services. So things like code were almost entirely skipped. He only mentioned code, but wouldn't be surprised if math etc was treated similarly. My assumption here is adding these missing subject matter areas had a larger impact than raw token counts.