5 ms·
That's Common Crawl, they do the spidering of some billions of webpages but that's still a tiny percentage of the web versus Google or Bing.
by web007 5y ago
That's Common Crawl, they do the spidering of some billions of webpages but that's still a tiny percentage of the web versus Google or Bing.
- cschmidt 5y agoDo you have any stats on that? I've always wondered about the coverage of Common Crawl, if you include all the historical crawl files too.
- visarga 5y agoCommon Crawl is being used to train the likes of GPT-3 and mine image-text pairs for CLIP. I wonder how much useful content is missing, we're going to use all the web text, images and video soon and then what do we do? We run out of natural content. No more scaling laws.
- collin128 5y agoOh interesting, I've played with it a little but not a dev and I've always wondered what the coverage was like.