5 ms·
Building a search engine to rival Google
- phenkdo 6y agoI think it's quixotic to compete with Google/Bing on the index size. A good bet would be to try to compete on search flexibility i.e. allowing for various ways of slicing/dicing results, recombinance etc. IMHO Search needs a paradigm shift from the 10 blue links to a knowledge engine (a la Wolfram alpha). The vast majority of searches can be satisfied by some of the open source indexes. For the "long tail", there is always Google.
- uniqueid 6y agoI agree. I don't know anyone my age who notices an improvement in the ability to find online information since about 2005. Yet the quantity of pages on the internet has grown by an order or magnitude over that time. Good search is no longer a matter of indexing everything; it's a matter of indexing enough high quality information.
- phenkdo 6y ago>Good search is no longer a matter of indexing everything; it's a matter of indexing enough high quality information. and I would add ability to digest that information into something usable would be next.
- bogomipz 6y agoCan you or someone else recommend any resources on the knowledge engine/"Wolfram alpha" architecture and how it differs from a traditional inverted index style search engine?
- bogomipz 6y agoThe article states: >"In its submission to the UK competition inquiry, Microsoft said that websites prioritize access to Google's web crawler over Bing's. As a result, Google is able to index these pages more thoroughly and deeply." Could someone say how this crawling priority works exactly? What is the thing being prioritized, the number of concurrent crawls by their crawlers? A different robots.txt for only the google domain? Something else?
- jsnell 6y agoRobots.txt has a directive for crawl delay, and like all directives it can be applied to specific user agents. In addition to that, the webmaster console can be used to set the crawl speed, and that data would obviously not be readable to third parties.
- waynesonfire 6y ago> The key to Google's dominance is not the fancy software or the ranking algorithms, but the sheer size of the index, said Zack Maril, a Washington DC software engineer who recently briefed US government investigators on Google's search monopoly. Why is the ranking not that important? Index all you want but if the ranking is done poorly that's a bad user experience. Or is the state of the art in information retrieval good enough and thus, doesn't cost much?
- zinekeller 6y ago> Or is the state of the art in information retrieval good enough and thus, doesn't cost much? Yes. Of course, you need to find the correct bias so that a satisfactory result is perceived by the greatest amount of people possible, but otherwise the real challenge is to convince websites to allow them a competition besides Bing and Google (which are often the only whitelisted "bots" for websites outside Russia and east Asia). For example, DDG has tried to use "Orange" (the name of their bot) previously in exchange to Bing but it didn't really work as expected due to websites whitelisting only MSN/Bing and Google. Google, LinkedIn and some other sites blocks access to their site index, giving boost to only established spiders. Even Twitter blocks Bing [1] as we speak, therefore only Google (and Yandex and Yahoo! if their indexer is still alive) can reliably parse Tweets. 1: https://twitter.com/robots.txt https://twitter.com/robots.txt
- strikelaserclaw 6y agoIt is highly concerning to me that tech is being sliced up by a couple key players and those key players are becoming more and more entrenched. Mobile apps/mobile web is how most of the world spends most of its time interacting on the web, and we only have 2 real options as far as mobile platforms go ios and android (Apple, Google). Search is dominated by Google. Browser is dominated by google. Social media is dominated by facebook. AWS powers god knows how much of the web. I wonder if as a collective effort open source community could somehow challenge one of these giants in their respective domains and provide alternatives.
- throwaway2016a 6y agoI always wondered if a distributed trustless search engine is possible. There are a few out there but non-of them come close in terms of result quality or speed. You also run into an issue where if the nodes are doing the indexing they can influence how much priority results have but if you can overcome that it would be very interesting. Content of the article aside... As an ex-Lycos employee I find it so strange how it seems to be wiped out of peoples' memory. Lycos was an innovator and top contender at one point. That image of search engines has several companies Lycos bought but not Lycos itself. I'd say it is because Lycos predates Google but Altavista is on there and others.
- jasfi 6y agoI tried working on a distributed search engine years ago (around 2006). It's very difficult to integrate results from different nodes with a global ranking that makes sense, at least from what I remember. I remember Lycos, but Google left all other contenders far behind, and most people only remember the winner.