5 ms·
Seems like it when for 7b models: (but also wow things are moving fast!); Llama - 1.4 Trillion tokens - feb. 2023, Llama 2 - 2 Trillion tokens - july. 2023,
by MyFirstSass 2y ago
Seems like it when for 7b models: (but also wow things are moving fast!);
Llama - 1.4 Trillion tokens - feb. 2023,
Llama 2 - 2 Trillion tokens - july. 2023,
Mistral - 8 Trillion tokens - sep. 2023 (this was the first big impressive leap where local models really became useful for chat)
Llama 3 - 15 Trillion tokens
So we went 1.4x, to 4x, to 2x. I wonder if there's even unused data out there?
Haven't tried the new model enough to see how much better it is than Mistral, that will be the real SOTA test for now.
- Escapado 2y agoI wondered about the same thing but at the same time only about 5% of the training data is non-English and I would be surprised if the total amount of published text from all non English languages combined was also just 5% of all published text. So my intuition tells me there is still heaps of data but what might be tricky is to properly access and asses it’s quality. Also the redpyjama v2 dataset has 30T tokens and is based on common crawl. Now I don’t know too much about common crawl but I doubt it has in it all published scientific books in all the different languages as these are often not freely crawlable. I remember when I studied physics there were at least 20 different 200-800 page long books on particle physics in German alone in our campus library. That must amount to 5million token by itself, from just one niche of physics. The Hamburg public library hosts about 5 million books and 90 million scientific articles mostly in English and German. If the average length of a scientific article is 6000 tokens and the average book about 100000 then that alone is already 1 trillion token. I bet there are significantly larger libraries and this is before even crawling the internet and looking at other languages or even generating training data.
- sheepscreek 2y agoStill quite impressive to think these models are trained on 15x the content in a large public library. That is insane.
- vidarh 2y agoDeutsche Nationalbibliothek appears to have 43.2 million "items", of which apparently 17.3 million are books. If we assume ~60,000 tokens for an average book, which seems very conservative given average word length in German and a novel typically being considered anything above ~40k words), that's another trillion just for their books, so I'm guessing the total German language content available in major libraries will be many times that. E.g. the Norwegian National Library has somewhere between 3x and 10x as many tokens in Norwegian newspapers as in books (at one point I think GPT3 breakdown of training data by language surfaced, and the Norwegian data was a tiny fraction of what was available in the national library, even before trying to estimate online/digital content). While I'm sure there's overlap [1] between the languages, a lot of it will help translation, and I think even for smaller languages the ratio of local content seems to dwarf translations. E.g. the "bestsellers" from English, French, and German all get translated to Norwegian, but most of the "long tail" content is local. [1] I was tickled to a find one of my uncles represented in Deutsche Nationalbibliothek; he was a professor in statistics, so it was a translation of some of his research
- londons_explore 2y ago> I wonder if there's even unused data out there? One day someone is going to train on the contents of DM's/private conversations/emails. There has to be 50x or more the quantity of that compared to public text. I suspect they'll do it via some 'prove-ably private training' regime, and therefore be able to claim it isn't a privacy violation.
- nolok 2y agoYes, they're either already getting into it behind legal facade, or aiming for it as the next eldorado of data. Facebook has whatsapp and messenger, microsoft has skype and outlook and exchange and msn messenger, google has gmail and all your text message and their bazillion chat apps and usenet and irc and ..., apple has imessage and icloud emails and ... There is so much data there, it dwarfs those token count.
- londons_explore 2y agoThe players with e2e encryption (imessage, whatsapp) would need to do some kind of client side edge device training. Possible, but hard to do with nobody knowing, and edge device training usually involves big quality compromises.
- nolok 2y agoDidn't both of them have some of their backup in clear text ? I know whatsapp backup on gmail were.
- Workaccount2 2y agoGoogle is sitting on ~20 years of gmail, but I can imagine the headache of both cleaning the dataset and likely consumer blow back. They also have youtube, which almost certainly has enough good data to train a powerful model on it's own, but also seems daunting to clean up first.
- seunosewa 2y ago
- nolok 2y ago> I wonder if there's even unused data out there? I think you're massively underestimating the amount of data out there. The challenge is how to access and categorize that data. Every usenet message, forum post from old school bbs to php forums to modern javascript abomination and closed discord boards, every email, every text message, every irc message, ... Those are probably a pain point to access due to rules and regulation and yada yada, but that alone dwarfs the 15 trillions, and you've not even started on actual quality content. (not saying these would specifically be good for llm, just answering to your "unused data" assessment)
- MyFirstSass 2y agoThat's a good point, though i already thought the OpenAI team had been very agressive in sweeping both reddit, usenet, + various illegal megatorrents of books, forum dumps etc. I remember there were some controversy around it on twitter a few months ago. One thing though is books/content/media from other language spheres though that could probably at least 10x the size of the data, and as far as i know translation starts to work rather well in these larger models so it would probably just plug right into the knowledgegraph for all languages?
- vidarh 2y agoThere are still vast amounts of data locked up behind login screens etc., though. E.g. to the foreign language data, a lot of national libraries around the world are either not even fully digitized yet or have lots of locked-down content. The Norwegian one is pretty open, but there's still huge amounts (like most newspapers newer than a century or so) that is either only available based on geolocation (I have my VPN for genealogy because of that - I'm Norwegian but live in the UK, and it's a nuisance), or only in a physical library in Norway. Similarly I was looking for something from the British Library at one point and it was behind a paywall (a copying fee). I have no idea how to even start to estimate how much data is locked down like that, and it's harder yet to try to figure out which parts of that it'd be possible to negotiate access to for various players, and what they can circumvent (e.g. say by buying book collections and the like - OpenAI is large enough by market cap it could afford to buy some of the largest extant publishers, for example, if they thought it gave them sufficient benefits).