5 ms·
RedPajama v2 Open Dataset with 30T Tokens for Training LLMs
- timcobb 3y agoSuper cool people are doing this. But I wonder: how will training data be any different from password lists of yore, which were the arms race secret sauce that no one ever shared?
- tydunn 3y agoThis is a lot of tokens. Llama 2 was trained on two trillion tokens [1] [1] https://arxiv.org/abs/2307.09288 https://arxiv.org/abs/2307.09288
- amilios 3y agoLoss was still decreasing for the models, there's a sense that we can push the training data much much further than we currently are.
- npsomaratna 3y agoYup. I found this article quite enlightening: https://espadrine.github.io/blog/posts/chinchilla-s-death.html https://espadrine.github.io/blog/posts/chinchilla-s-death.ht...
- rushingcreek 3y agoPhenomenal blog post about scaling laws.
- deleted 3y ago[deleted]
- famouswaffles 3y agoPrediction as an objective basically forces the models to model the casual processes that create the text itself. It's not going to stop getting better unless the data is insufficient/unvaried or the architecture creates a bottleneck. I think by the time the former is an "issue", we'll have a Super Intelligence on our hands anyway. The latter is looking less and less likely to be a real hurdle. Very little inductive bias to steer away from crucial solutions, very scalable.
- yorwba 3y agoThe TinyLlama project is trying to do that pushing by training a small 1.1 billion-parameter model on 3 trillion tokens: https://github.com/jzhang38/TinyLlama https://github.com/jzhang38/TinyLlama
- minimaxir 3y agoAnd Llama 2's training data was less aggressively deduplicated.
- artninja1988 3y agoNice. Hope somebody makes a torrent of it/ hosts it in a way that it can't be taken down. Also, what are some estimates of how many tokens of text are out there? Seems like we are hitting that number pretty quick?
- civilitty 3y ago> Seems like we are hitting that number pretty quick? I don't think we're even close. Libgen's nonfiction archive alone is over 32 terabytes. Total size last year was over 120 terabytes. Between that, SciHub, and the internet, there's probably orders of magnitude more tokens out there.
- famouswaffles 3y agoI don't know about orders of magnitude left but we're definitely not close yet. This is just 5 languages(and frankly not even the 5 with the most text) and just as importantly, just what is crawlable from the web. There's tons of stuff in popular ebook archives you can't crawl from the web. This is also relatively code/scientific corpora scant. We're just getting started.
- jprete 3y agoIt looks like mass copyright infringement, frankly.
- rizky05 3y ago[dead]
- xvector 3y agoIf it gets us to AGI faster, I frankly don't give a fuck. AGI-driven drug discovery will save billions of lives. Every day it is delayed costs tens of thousands of lives. No amount of copyright is worth that sacrifice.
- artninja1988 3y agoHeh. Bit too high of a bar. Even if it helps to develop boo about 9000 that's fun to prooomt for a while, I think it's fair game
- peddling-brink 3y agoAGI will mean we’re no longer the dominant life form on the planet. If AGI were achieved tomorrow, how many humans would be left in 200 years?
- soultrees 3y agoI find this notion interesting. What makes you think AI will automatically kill humans?
- pr337h4m 3y agoAnd moreover, what makes people even think that a desire to commit mass murder is an innate characteristic of an 'intelligent' being, that increases the more 'intelligent' it becomes? (If they believe themselves to be 'intelligent', do they believe they have a greater desire to commit mass murder?)
- gardnr 3y agoAnyone know how large it is? They state the 1 trillion token dataset is 5TB. Is it safe to assume this is 5TB * 30 = 150TB? The code in the HuggingFace repo downloads data from url base: https://data.together.xyz/redpajama-data-v2/v1.0.0 https://data.together.xyz/redpajama-data-v2/v1.0.0 https://huggingface.co/datasets/togethercomputer/RedPajama-Data-V2/blob/main/RedPajama-Data-V2.py https://huggingface.co/datasets/togethercomputer/RedPajama-D...
- zhangce 3y agoIt is around 100TB (84 CommonCrawl dumps, roughly 1TB per dump)
- mauriceweber 3y agoyes, small clarification: the 1TB per dump refers to the head+middle partition of the dataset and includes the text documents and the quality signals. There is another ~700GB for the minhash signatures and 1-1.5TB for the documents in the tail split.
- natch 3y agoCan someone explain to me like a noob how this ("this" being the data hosting and download access) works? Am I understanding correctly that they are releasing code for filtering common crawl data that is out there, and the result of this filtering is the dataset? To further elaborate on this (possibly wrong) understanding: - Each person can then run their own processing, possibly duplicating effort(?) ...but on the good side, giving each person the ability to tweak the pipeline to suit their needs. - There is no torrent of already processed data because __________? - Looking at file lists for this on Hugging Face, some files seem to be stored in Git Large File Storage. Are these already processed files that together constitute the dataset? Or are these Common Crawl files that are selectively listed and pulled for processing? What options are there to preemptively obtain a copy, in case of any possible eventual takedown of the dataset, any assurances about access aside? I am reminded of parts of the pile. Obviously I'm super clueless here... please be gentle and share anything you know or correct anything I've got wrong. I'm not asking about training, if that wasn't obvious. Just about obtaining the dataset.
- zhangce 3y agoWhat we make available is: -- (A) the dataset after pre-processing the raw CommonCrawl data (e.g., text extraction and language identification) and some minimal filtering; and (B) for each document in (A), we also pre-computed 40+ of "features" (we call the "quality annotations") you can use to further filter it or deduplicate it. For example, one such feature is "how similar this document is to Wikipedia". -- (A) is around 30T tokens, but you might want to use features in (B) to further filter/dedup it down, e.g., to 5T. For example, if in your application documents similar to Wikipedia are the most helpful documents, you can take the top documents with the highest score for the feature "how similar this document is to Wikipedia". Of course, the really interesting case happens when you consider a larger subset of these features (or maybe even automatically learn what the best way of filtering it is). Our goal is to make this as flexible as possible such that you can fit this into your own application. What we have released is both (A) and (B) If you have any questions, please let us know! Thanks for your interests, have fun with the data!
- natch 3y ago
- applgo443 3y agoIf it's 5 common crawls, isn't data across multiple common crawls mostly similar?
- zhangce 3y agoWe did an exact dedup across all 84 dumps; there are 100T tokens before this exact dedup, and 30T tokens after. If we do further fuzzy dedup (we have simhash signatures pre-computed for different similarity level), this can potentially be reduced further. There are quite a lot redundancies across dumps; but also a lot of unique/distinct documents
- visarga 3y agoGreat work, may I suggest more analysis features? - example summary, for better topic embedding - RAG based summary, to have the model critically assess its training data distribution and answer questions on it; to bring together information sitting in separate examples - named entities, for knowledge base; maybe it helps with fact checking later - implicit tasks present in the text, what are the tasks a LLM could learn from a given example? - chain-of-thought augmentation, to bring out implicit deductions and reduce information fragmentation; it has been shown in the Phi-1.5 paper and Orca that synthetic CoT datasets are superior source materials What data fragmentation? Look at the Reversal Curse paper. Models that train on "A is the father of B" fail to generate "B is the son of A". This kind of connection needs to be explicitly added, and would improve task solving as well. Training on purely organic data is not good enough anymore. All powerful models train on a mix of organic and synthetic data, some models on 50-50 proportions, like the web+synth variant from Phi-1.5. The main idea is to go deeper into the raw data, to infuse it with insight. LLM dataset preprocessing is going to be expensive, comparable to training costs, but the results are worth the effort.
- zhangce 3y agoThanks for the suggestion! We will add this in the pool of features for future release. (We are currently running the current 40+ annotations on the `tail` partitions). If you are interested in contributing the code for these features, feel free to do a PR to https://github.com/togethercomputer/RedPajama-Data https://github.com/togethercomputer/RedPajama-Data! Otherwise we will try our best effort implementation :) but we hope that this can become a community effort (feel free to created more issues on github for us to keep track. I created one for this https://github.com/togethercomputer/RedPajama-Data/issues/76 https://github.com/togethercomputer/RedPajama-Data/issues/76)
- sorokod 3y ago"B is the son of A" doesn't follow from "A is the father of B". B could be A's daughter.
- TeMPOraL 3y agoI feel this is usually a hugely asymmetric problem. The other example I've seen is the model being able to easily complete "The color of the sky is" with "blue", but then failing to complete "Blue is the color of" with "the sky". I say, d'uh, why would it? If you take the training data, or in general, imagine taking all that humans ever wrote or spoke in English to date, you'll expect to find an overwhelming amount of cases where "The color of the sky is" ends with "blue". However, "Blue is the color of" can easily have a hundred thousand different plausible completions, and "the sky" won't even be one of the more likely ones. In the absence of additional context that strongly hints at the answer, one should NOT expect a properly working LLM to frequently propose "the sky" as completion to "Blue is the color of".
- shoelessone 3y agoThere are so many articles these days posted on HN like this recently but I'm realizing I am too far out of touch with the technology to be able to appreciate it. Any recommendations as to how I get a bit of hands on experience in the AI "domain" so when I read some news articles like this it means something more to me? Or is this type of thing really only relevant to a very small subset of software people?
- all2 3y agoThere's a course available here [0] that might interest you. [0] https://www.fast.ai https://www.fast.ai
- e12e 3y agoNice. I admit I find the language selection a bit uninspired and odd: > Five languages: English, French, Spanish, German, and Italian Otoh I'm surprised that when counting first and second language proficiency, German is actually ahead of Japanese.... https://en.m.wikipedia.org/wiki/List_of_languages_by_total_number_of_speakers https://en.m.wikipedia.org/wiki/List_of_languages_by_total_n...
- kiney 3y agoI'm surprise for the low number native german speakers in that list. Thats less than the population of germany alone. And then theres austria, parts of switzerlan, northern italy, western belgium....
- deepsquirrelnet 3y agoI’ve been impressed with “fuzzy” deduplication at this data scale. I’ve used minhash and networkx for small amounts of data, but I really appreciated the write up on your GitHub about how you implemented it for this dataset.