5 ms·
I've tried using different search engines to Google numerous times, but each time I've returned to Google simply because the searches are better. They're more a
by libeclipse 11y ago
I've tried using different search engines to Google numerous times, but each time I've returned to Google simply because the searches are better. They're more accurate, more relevant, and I very rarely find myself searching more than once to find something.
If commonsearch can beat Google in that regard, then count me in. But I doubt it will.
- JohnKacz 11y agoWhen I switched to DuckDuckGo last year I read an interesting comment from someone. The basic idea was that we have all become so accustomed to Google's results and the manner in which we use it (i.e. the way we define our search terms) that it is actually we who must be willing to reprogram our search practices if any competitor is to have a chance to catch up. I'm not sure if I buy that, but I do believe if we don't commit to alternatives it will be next to impossible for alternative search engines to get as good as Google with result quality and relevance. Google simply know too much about me and has performed so many more searches for a rival to outperform them. I still use DDG's 'g!' often but I feel like I'm doing my part to help DDG get better for me and other users.
- djsumdog 11y agoSame here. I say I'd use !g on duckduckgo about 1/3 of the time. If I'm getting really frustrated on a problem and am hacking away, I sometimes just default to the !g. Duckduckgo is now doing localized results (you can choose your region, so it's transparent, unlike Google's). The thing about Google is that, even if you're not signed in, it still tried to present you with personalized results (based on previous searches for that session, your IP, your region ... if you're searching from work; it probably factors that in as well). When people talk about getting to the top of Google results, my response has always been, "Well you need to be more popular and relevant. Also you may be at the top..for some people, but not everyone."
- sylvinus 11y agoI definitely share your usage of "!g", which is why Common Search already supports it ;)
- sylvinus 11y agoHi! I'm the founder of Common Search. I don't think search result quality is on a linear scale so it's hard to define "better". The results will definitely be less personalized, which will be a big plus for some people, and a blocker for others. There will be a few other dimensions where we can stand out, and some where we will have a hard time catching up (index size for instance). In the end, given enough contributors, I'm pretty sure the results can get "good enough" for most people, and hopefully "better" for some ;)
- mschoebel 11y agoDo you have a rough estimate of how many servers you will need for Elasticsearch for the 1.7bn URLs of the latest CommonCrawl?
- sylvinus 11y agoThe current average size of documents in the index is 2kB, so we'll need ~3.5TB of storage. For 1 replica this could mean 5 i2.xlarge instances if we go for SSDs on AWS.
- mschoebel 11y agoThanks for the answer. I suggest not to use AWS if you know that you'll need a server 24/7. Old-school hosters which offer dedicated servers are much cheaper for that use-case. There are several offers here in Europe where you can get an i7-6700, 64gb RAM and 1tb SSD for less than €60/month. AWS would cost you at least 3-4x as much. You'll lose the flexibility of AWS, but save a ton of cash.
- jasode 11y ago>AWS would cost you at least 3-4x as much. You'll lose the flexibility of AWS, but save a ton of cash. Isn't there more to the analysis than just comparing cpu before we can conclude it will save a lot of money? It looks like their servers[1] use ~150TB source data that's already hosted on AWS disks. The source x.gz archives of the Common Crawl on AWS S3 are then imported to a Elasticsearch disks that are hosted on AWS. To pull ~150TB of data using network speeds of 30 megabytes/sec[2] would take 60 days to transfer from AWS to another USA datacenter like Rackspace. (Copying data from AWS to AWS isn't instantaneous either but it won't take ~60 days. At 60 days, the next crawl archive would have been released before you finished importing the previous one!) Questions would be: 1) What are current 2016 network speeds between cloud providers? 2) What's the cost of ~150TB of network bandwidth? 3) From those datapoints, can we derive a rough rule-of-thumb where a certain amount of data exceeds the current capabilities (speed or economics) of the internet backbone available to projects like Common Search? [1]https://about.commonsearch.org/developer/operations https://about.commonsearch.org/developer/operations [2]http://www.networkworld.com/article/2187021/cloud-computing/amazon-comes-out-on-top-in-cloud-data-transfer-speed-test.html http://www.networkworld.com/article/2187021/cloud-computing/...
- nostalgiac 11y agoI agree. The times I search on computers in the address bar and then wonder why the results are bafflingly awful, I realise the search is set to Bing or any other Search Engine. Re-typing the exact same phrase into Google and suddenly the top results are exactly what I wanted in the first place.
- erik14th 11y agoI believe trying to create a general search that's better than google is pretty much a lost battle. I think the next big thing in search will start as a niche thing. If you reduce your user domain you can have a better shot at producing better results even without tracking. For example, you could develop a search engine for developers/IT people, make it's results better than google's and then expand to other domains.
- zookatron 11y agoI think what we really need is basically exactly what Google has in terms of search technology, except that it needs to be open and explicit, instead of it being Google's secret proprietary data on me that I can never access and that they get to sell to advertisers to my detriment. I (like many other people on HN) use Google constantly when programming, and it's impossible to overstate the convenience and power of Google's almost creepy ability to guess exactly what the language and context of my search query is. It cuts precious seconds off of each query (when I am routinely make hundreds of queries per day), and more importantly cuts out the interruption of mental flow as you try to re-word your query into a format that the search engine will understand. This can potentially add up to hours of saved time per day, depending on how you calculate the impact of these features. In order to duplicate this I don't think we can get around the need for "search profiles", which takes into account your location, interests, past searches, personal connections, etc, but it needs to be explicit and it needs to be my data. If I want to delete it or sell it, it really needs to be up to me. If we could figure out a secure way to do this, then we would have the framework necessary to compete with Google with an open platform. Until then, it's just not going to be worth it to switch for the vast majority of people.