4 ms·
Have you ever looked at the Amazon file? I'll see if I can track down the link but I remember somebody sharing a dump with me from Amazon that apparently was a
by collin128 5y ago
Have you ever looked at the Amazon file?
I'll see if I can track down the link but I remember somebody sharing a dump with me from Amazon that apparently was a recent scrape.
Edit: https://registry.opendata.aws/commoncrawl/ https://registry.opendata.aws/commoncrawl/
- web007 5y agoThat's Common Crawl, they do the spidering of some billions of webpages but that's still a tiny percentage of the web versus Google or Bing.
- cschmidt 5y agoDo you have any stats on that? I've always wondered about the coverage of Common Crawl, if you include all the historical crawl files too.
- visarga 5y agoCommon Crawl is being used to train the likes of GPT-3 and mine image-text pairs for CLIP. I wonder how much useful content is missing, we're going to use all the web text, images and video soon and then what do we do? We run out of natural content. No more scaling laws.
- collin128 5y agoOh interesting, I've played with it a little but not a dev and I've always wondered what the coverage was like.