4 ms·
I'm not sure splitting off the crawler would have the effect you want it to. Crawling is hard, but actually one of the easier parts of building a search engine.
by GradientAssent 6y ago
I'm not sure splitting off the crawler would have the effect you want it to. Crawling is hard, but actually one of the easier parts of building a search engine. There are less complete, but still pretty impressive public crawls available already, like Common Crawl.
The harder problem is what you do with the result of the crawl once you have it. You need to index it. This is one of Google's deepest moats. Being able to serve any of hundreds of billions of documents in milliseconds requires an enormous amount of infrastructure and hundreds of careers worth of difficult software engineering.
I don't see a good way to split indexing out as a separate company from the rest of the search engine. It's tightly integrated with ranking. Generally a new ranking technique requires some new information to be associated with each document in the index, e.g. some precomputed scores that describe the document's quality or fitness for different kinds of queries. At query time, you're often just weighting and comparing these scores across documents.
If you had Google's index, you'd be a long way towards being able to replicate Google's results, but I don't think that would lead to the kind of innovation you're hoping for.
- absolutelyrad 6y ago> Common Crawl Not even close. I've used Common Crawl, it has nowhere near the breadth that Google has in their index. Problem no. 1 is that website owners block non Google crawlers because of limited bandwidth. So you can't reliably even get started on creating your own search engine. Once we have the data available, we can figure out the other bits and what they require. But to get to that door, you need access to the crawl. You don't need to have everything in the Google index to beat Google. You need to be better than Google in the search niche that you'll serve. And for that, at the bare minimum, you need access to all the crawl data that Google has to get started, then throw away the bits that you don't need.
- oxygenjoe 6y agoSo if we make internet bandwidth a regulated utility like electricity and water, would that create an environment where other crawlers could compete?
- deleted 6y ago[deleted]
- wbl 6y agoBandwidth costs have been continually declining for decades since competition started. The utility model of Ma Bell kept bandwidth costs up.
- technotarek 6y agoWell, I wasn't taking the OP's idea to mean only splitting off the crawler. Why not an API where you send it the search query and then it returns the results? It could have a variety of return options including using a ranking of the API publisher's choosing or you could ask for something raw, and innovate with an even better or more suitable ranking based on your specific needs. In the end, the point would be to make Google closer to what it was at the beginning -- just input and output. No ads on top of it, no bundling with other offerings etc. And if the publisher's ranking remains as good as it should, then surely some will be willing to pay not only for the raw data, but the ranking as well.
- visarga 6y agoThis is exactly how I envisioned it. A crawl+index company that gives equal access to all competitors. The search front companies would use the index to retrieve pages, then they would rank these pages with their own criteria and filter with their own white/black lists, then present the result in their own UI. Each one would have their own ads and tracking (or lack thereof). The index would not be able to know the identities of the users. In the end they all boil down to an "attention economy", it's about us and our eyeballs, we should have more say on what we are consuming. The part about ranking, filtering, tracking and UI is true for FaceBook and Twitter as well, even though they are a different kind of service. We could have alternative views on these systems, where we have more control on the way the system manipulates the information for us.
- dmitriid 6y agoGoogle's crawler gets preferential treatment on many sites. Not necessarily because of private deals, but because it drives traffic from search results. Quote via [2] --- start quote --- When Mr. Maril started researching how sites treated Google’s crawler, he downloaded 17 million so-called robots.txt files — essentially rules of the road posted by nearly every website laying out where crawlers can go — and found many examples where Google had greater access than competitors. ScienceDirect, a site for peer-reviewed papers, permits only Google’s crawler to have access to links containing PDF documents. Only Google’s computers get access to listings on PBS Kids. On Alibaba.com, the U.S. site of the Chinese e-commerce giant Alibaba, only Google’s crawler is given access to pages that list products. --- end quote --- [1] https://www.nytimes.com/2020/12/14/technology/how-google-dominates.html https://www.nytimes.com/2020/12/14/technology/how-google-dom... [2] https://daringfireball.net/linked/2020/12/14/wakabayashi-google-search https://daringfireball.net/linked/2020/12/14/wakabayashi-goo...
- sib 6y agoSo it sounds like the regulators should focus on forcing every web property to give equal access to all crawlers. Why steal Google's investment and hard work and innovation?
- hamilyon2 6y agoThere is so much irony in the fact that the web loves Google crawler, optimizes for it, makes sure that Google robots have special, first class treatment. And at the same time, on the same website I need to enter Google-provided captcha. To fight robots.
- coldtea 6y ago>I'm not sure splitting off the crawler would have the effect you want it to. Crawling is hard, but actually one of the easier parts of building a search engine. There are less complete, but still pretty impressive public crawls available already, like Common Crawl. When they saw crawler they don't mean Googlebot or something. They mean the whole thing you've described, the search engine as a backend. What the law proposes is that other companies should be able to just buy the search engine service from a new company split from Google, of which Google which will just be just another client. >I don't see a good way to split indexing out as a separate company from the rest of the search engine. It's tightly integrated with ranking. The idea is to split the whole search part from Google. Let them have the ad business, Android and so on, and just rent the search from the spin-off company...