5 ms·
It's not that easy. Access to enough compute is one thing. However, you also need a proper dataset (beyond Common Crawl and Wikipedia), excellent research exper
by beernet 4y ago
It's not that easy. Access to enough compute is one thing. However, you also need a proper dataset (beyond Common Crawl and Wikipedia), excellent research expertise and engineering capabilities. So even if you throw money or free credits for cloud compute out there it will not be enough. We've seen this happen with EleutherAI who were not capable of reaching their initial target of "replicating" GPT-3 and could only deliver the GPT-NeoX 20B model despite all the free compute etc.
- rjh29 4y agoThere is Open Street Map (or Wikipedia for that matter). A large enough army of volunteers could produce or tag a dataset that rivals Google's data, but it would be a lot of work.
- orbz 4y agoAgreed, I was thinking the same thing. Raw compute is probably the cheapest/easiest part of this problem.
- oliwary 4y agoThis is my impression as well. Here is an example of the engineering capabilities required to train the modern models: https://github.com/facebookresearch/metaseq/blob/main/projects/OPT/chronicles/OPT175B_Logbook.pdf https://github.com/facebookresearch/metaseq/blob/main/projec... [PDF] It's a 114 page document detailing the months of full-time work for multiple engineers that went into training a 175B language model at meta AI.
- sillysaurusx 4y agoWe solved the proper dataset part at least. https://arxiv.org/abs/2101.00027 https://arxiv.org/abs/2101.00027 My contribution was around 19,000 books.
- hooande 4y agoisn't Common Crawl much, much larger than this? ~6 pebibytes from what I remember
- sillysaurusx 4y agoYeah, but the hard part is filtering. It’s pretty easy to scrape a massive amount. Turning it into quality training data is the trick. bmk was the magician there. https://twitter.com/nabla_theta?s=21&t=Gt6YrATJHnmY046MdzhYDg https://twitter.com/nabla_theta?s=21&t=Gt6YrATJHnmY046MdzhYD...
- swyx 4y ago> My contribution was around 19,000 books what does this mean? not meaning to cross examine you, just curious how people contribute to The Pile since it seemingly appeared out of nowhere
- sillysaurusx 4y agoNot at all, I love talking about it. I was convinced that a model needed to be able to read like we do. And what do we do when we read? Pick up a book. That turns out to be surprisingly hard, at least for training data. Step one is to acquire the books. Step two is to turn them into a readable format for computers. Both steps were very hard. I lucked out on step one because The Eye happened to host all of bibliotok, which came to around 30k books or so. Trouble is, lots of those are PDFs. And although humans are great at reading those, they fucking suck for blind people. And a gpt is a blind person in a sense, because it needs to follow a linear sequence of words — something that PDFs are horrible at giving. But one day I realized that epubs were merely html files, and aaronsw happened to write an amazing html to text converter. I had to hack it to fix a few corner cases. But after a few days, I ran it across all 19,000 epubs I spidered, then zipped the whole thing up and called it books3: https://twitter.com/theshawwn/status/1320282149329784833?s=46&t=mecILIS2Eh73ASUY2N95YQ https://twitter.com/theshawwn/status/1320282149329784833?s=4... It’s one of the larger components of the pile, I think around 35%. Which is quite the hefty sum when it’s purely text. I still have a hard time wrapping my head around just how mindbogglingly big 800GB of text is.
- sinenomine 4y agoData isn't the hard part here, plenty is available, even with all the necessary preprocessing.
- lossolo 4y agoA large amount of data is not an issue, but obtaining a high amount of high-quality data is challenging. This is why open-source models do not perform as well as GPT-3 models in real-world usage.
- sinenomine 4y agoSorry, since the Pile and C4 https://huggingface.co/datasets/c4 https://huggingface.co/datasets/c4 (and, more generally, common crawl) and BigCode https://www.bigcode-project.org/ https://www.bigcode-project.org/ became available, this argument ceased being the real moat. The real moat is more about the lack of concentrated compute, ML engineering, and, more generally, prosaic lack of political with outside a few orgs.
- swyx 4y ago> Common Crawl is there a crowdsourced list of text corpuses somewhere? i bet thats the starting point for all this. i'm only aware of C4 and The Pile.