3 ms·
If someone would really want to break the Google search monopoly, they would seperate: - Crawling and document parsing. Imagine something like archive.org wher
by bluelu 6y ago
If someone would really want to break the Google search monopoly, they would seperate:
- Crawling and document parsing. Imagine something like archive.org where all search engines would get the content from. This could also be used to enforce copyright legislation (e.g. deletions of documents, etc...). Today, it's nearly impossible to collect the data as google does, because most website owners won't even allow that you crawl them (e.g. distill networks anti bot protection, etc...)
- Metrics (e.g. data from google analytics, dns resolvers, browser toolbards, browser statistics, etc...). Again this data should be available to everyone from a central authority.
Then there could be more independent search engines building on top of that data. If that doesn't happen, then there will never be any competition. Todays copyright and data privacy laws even further strengthen a monopoly position, as smaller players will not even get access to some data (or at different pricing then google does)
- ricardo81 6y agoHaving one crawler would not be desirable as it'd involve choosing which pages get crawled and how often for freshness. Point taken about sites being less available to crawl due to the bot whitelist approach but it is not a huge issue, though some large sites like Facebook go for a whitelist only approach. Regarding those other data sources, the more privacy conscious engines would perhaps object to particular kinds of data but agreed that the kind of scale and various sources that Google has at its disposal, does give it an advantage when it comes to search. There are alternatives beyond Google and Bing who do index the web albeit at the moment are smaller in scale, and likely outnumbered by all the Bing clones. There is a high barrier to entry wrt cost but the technical/capital hurdles of crawling are not so high as compared to ranking and indexing.