9 ms·
The Pile: An 800GB dataset of diverse text for language modeling (2020)
- Der_Einzige 3y agoI came so close to getting my dataset DebateSum (https://huggingface.co/datasets/Hellisotherpeople/DebateSum https://huggingface.co/datasets/Hellisotherpeople/DebateSum) into the pile, but they decided at the last minute not to add it: https://github.com/EleutherAI/the-pile/issues/56 https://github.com/EleutherAI/the-pile/issues/56 I'm still a tiny bit salty about that, but the pile is a wonderful dataset regardless.
- orange_fritter 3y agoThat dataset looks cool. Good work either way, I'm sure it'll go somewhere
- Der_Einzige 3y agoStay tuned! I've got a paper I'm writing about a new followup which is a 40x improvement in size (basically every open source debate card... Ever) and a 40x improvement in metadata and duplication detection. The work is all done since late april and I've just been lazy/writer-blocked (ironic in a world of high end LLMs) and haven't gotten the paper finished. Kinda of sad to have missed NeurIPS dataset track deadline and ACL, but I know that anything close to this in scope is a slam-dunk accept at the argument mining workshop
- robmsmt 3y agoWould love to see an early version of it!
- sillysaurusx 3y agoAuthor here. And by author I mean I created books3 (the books component of The Pile) while everyone else did the hard work of actually writing the paper, ha. Stella and Leo Gao in particular did so much wonderful work on the paper, though it couldn’t have happened without everyone’s contributions. As far as I know, this was the first academic contribution from a discord collaboration to ML. Back then discord was barely used for ML at all, though nowadays of course the largest discord in the world is midjourney. There were a bunch of interesting stories from those days. We almost didn’t release at all (or at least the books component) because of fear of copyright backlash. Turns out no one cared, and then suddenly today the world cares a great deal. As a side note, I’ll be participating in a legal action against Meta for the purpose of making ML models uncopyrightable: https://twitter.com/theshawwn/status/1641804013791215619?s=61&t=jQbmCk1JqL7depzFWJNuPA https://twitter.com/theshawwn/status/1641804013791215619?s=6.... They DMCA’ed one of my repos distributing LLaMA, so we fought back and challenged the idea that weights can be copyrighted at all. This seems like the best outcome for hackers and individual researchers, for a few reasons. It’s also one of the most ethical outcomes; since ~no one trains on data that they own, they shouldn’t own the resulting model. One last thing. The Pile would’ve been far less relevant without the wonderful assistance of The Eye, a group of people who archive all kinds of things. They’ve hosted the datasets for years now. And although it seems strange to say that dataset hosting could make or break The Pile, back then there was nobody else willing to host us. https://the-eye.eu/ https://the-eye.eu/
- sfriedr 3y agoCould you share more about copyright? For example, aren't you worried that now, with all kinds of lawsuits happening [1] and copyright issues that were found in existing datasets [2], that you might get threatening letters from a lawyer some day? I'm the author of [3] where we introduced one of the first natural-language datasets that test graduate mathematics for LLMs, but some of the prompts we took from a copyrighted book and therefore thought about excluding them. Having them in the public dataset would be really nice though, hence I'm keen about your experience. I'd also be keen to hear how your challenge against the DMCA on sharing LLaMA's weights goes? [1] https://www.theguardian.com/books/2023/jul/05/authors-file-a-lawsuit-against-openai-for-unlawfully-ingesting-their-books https://www.theguardian.com/books/2023/jul/05/authors-file-a... [2] https://arxiv.org/abs/2105.05241 https://arxiv.org/abs/2105.05241 [3] https://arxiv.org/abs/2301.13867 https://arxiv.org/abs/2301.13867
- Der_Einzige 3y agoGetting sued is straight up a good thing for most peoples careers in tech. Haven't you watched silicon valley?
- sillysaurusx 3y agoI think a lot of hackers shy away from doing impactful work because of fear. Sometimes those fears are justified, but it's remarkable how often things that seem like a big deal turn out not to matter. My advice for ambitious devs would be to do what seems interesting, and don't worry too much about threatening letters. Usually the worst thing that happens is that you agree to stop doing whatever generated the threat. Personally, I'm not worried. It would be a damn shame if academics come under fire merely for trying to operate on the cutting edge of science. None of us were trying to make money; we just wanted to make something interesting. > I'd also be keen to hear how your challenge against the DMCA on sharing LLaMA's weights goes? Thanks! I think we might be putting up a website for it soon, if only to explain ourselves. In the meantime – I hate this phrase, since I don't want followers – the only way to keep informed is to follow my Twitter, and perhaps keep an eye on my HN comments. You'll probably hear about it either way though, since it's a groundbreaking case. No one has tested the copyrightability of ML models before.
- cschmidt 3y agoIf you’re looking at The Pile, you also might consider the Red Pajama dataset. A new cleaned version was released recently https://www.cerebras.net/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of-redpajama https://www.cerebras.net/blog/slimpajama-a-627b-token-cleane...
- CamperBob2 3y agoIs there a straightforward way to download that dataset, the way there was for the original RedPajama data? SlimPajama appears to have been released as 60,000 small files, which is ridiculous.
- dang 3y agoRelated: The Pile: An 800GB Dataset of Diverse Text for Language Modeling - https://news.ycombinator.com/item?id=36272365 https://news.ycombinator.com/item?id=36272365 - June 2023 (5 comments) The Pile: An 800GB Dataset of Diverse Text for Language Modeling - https://news.ycombinator.com/item?id=25607809 https://news.ycombinator.com/item?id=25607809 - Jan 2021 (60 comments)
- charlysl 3y agoOP here. I learned about this while reading Stanford's LLM course's "Data" lecture [1]. Very interesting how it assesses the datasets used for GPT 2 and 3, etc, and how The Pile addresses their issues. A very interesting course! [1] https://stanford-cs324.github.io/winter2022/lectures/data/ https://stanford-cs324.github.io/winter2022/lectures/data/
- pjot 3y agoThe Pile was also referenced in a post today of some guys tweets about “leaked” gpt4 details https://news.ycombinator.com/item?id=36675934 https://news.ycombinator.com/item?id=36675934
- robertheadley 3y agoAs long as LLMs and generative AI uses copywritten works for training, then they are going to be the enemy of creative people.
- splatzone 3y agoUnless the financial benefit could be shared with the original authors somehow, with some kind of royalties system?
- fsckboy 3y agoI love how "creatives" enjoy the freedom of the free internet but never try to shame their peers as to whether they use GPL or MIT license for their art.
- robertheadley 3y agoI think the more matter of fact the influence, the more the original artist deserves compensation. See Waits v. Frito-Lay, Inc. I do not not want something like this to happen to generative AI and make things more difficult for the technology to progress and flourish.
- mattkevan 3y agoCreative people will be using LLMs and other models as new and exciting creative tools. Their real enemies will be the people who make money off the creative people’s work, e.g. the entire history of recorded music or the current writers strike.
- koheripbal 3y agoThis is like saying that my brain violates copyright when I write sci-fi because one time, years ago, I watched Star Wars.
- Roark66 3y agoGreat stuff, I skimmed the article searching for some table showing a breakdown of content by language, but I haven't found one. I hope there is a lot of text in languages other than English. As for example in my language (Polish) current SOTA models are very deffiecient. I have wondered why is that considering companies like (not at all)OpenAI claim to train on large datasets including in my language of interest. It turns out (and I learned this just yesterday) they used LLM translated English content that that used as other language training data. They used Azure translator which itself is a transformer model to generate content for gpt-3.5 for example. Also, I bet there is a lot of poorly machine translated content in their supposedly "original" data. The result? You can use chatgpt to write you an email of any kind in English and you can copy/paste/send immediately. Try doing that in Polish... It will make sense, but the language used will use bad tone (too familiar in a business setting), bad words(words that exist, but no real person would use) and sentence layout that just plainly feels weird. I suspect this is even worse in many other languages.
- koheripbal 3y agoWhile having multiple languages makes a model more versatile and appeal to a wider audience, it actually significantly increases the memory required to run the model and thus limits other aspects of the model. Optimally, a Polish audience should try to create a Polish trained model. As it stands now, most advanced models, like gpt are multilingual, but are noticeably less capable in non-English languages.
- spi 3y agoHaving every model re-trained in each language is a certain path towards having any non-English (or at most a couple of other languages from countries with big pockets, like Chinese) language model be always massively behind - the resources required to train a model are huge, you can't expect e.g. the Polish community (plus anyone else) to replicate every good English model that comes out. GPT4 is less capable in Polish than in English, but probably much more than any Polish-specific model ever trained - and I suspect the gap is bigger than that with the best non-GPT4 English model. Furthermore, I think you are exaggerating the memory issue of multilingual models significantly. Especially for languages using the same (Latin) script, the additional characters to care about are very few. Also a significant part of the vocabulary and language fall into a few buckets, so training a joint model makes all the sense in the world - much like an Italian native speaker could likely study a scientific text in Spanish and understand its content, even without speaking the language. The memory impact comes mostly from having bigger embedding layers that have to account for vocabulary in many languages (the most problematic case being Chinese and Japanese, with their huge set of tokens). But even there, the largest vocabularies in use are maybe of size 100k (vs. about 30k for English-only), with a hidden dimension of 4k that makes for a total of 400M parameters. It's a lot, but a drop in the ocean of 100B+ parameters (or 1T+ for GPT4) we're seeing today. P.S. Answering to GP, I think the Pile is English only, though - or at least, models on HuggingFace trained on the Pile, like the various Pythia models, are tagged as English only.
- ryoshiro 3y agoSide Topic: In the leaked OpenAI GPT-training details, there are speculations that OpenAI trained on Libgen dataset. Is there a link to the dataset of Libgen, if so how big is it?