3 ms·
> Common Crawl Not even close. I've used Common Crawl, it has nowhere near the breadth that Google has in their index. Problem no. 1 is that website owners bl
by absolutelyrad 6y ago
> Common Crawl
Not even close. I've used Common Crawl, it has nowhere near the breadth that Google has in their index.
Problem no. 1 is that website owners block non Google crawlers because of limited bandwidth. So you can't reliably even get started on creating your own search engine. Once we have the data available, we can figure out the other bits and what they require. But to get to that door, you need access to the crawl.
You don't need to have everything in the Google index to beat Google.
You need to be better than Google in the search niche that you'll serve.
And for that, at the bare minimum, you need access to all the crawl data that Google has to get started, then throw away the bits that you don't need.
- oxygenjoe 6y agoSo if we make internet bandwidth a regulated utility like electricity and water, would that create an environment where other crawlers could compete?
- deleted 6y ago[deleted]
- wbl 6y agoBandwidth costs have been continually declining for decades since competition started. The utility model of Ma Bell kept bandwidth costs up.