8 ms·
Common Search – nonprofit search engine for the Web
- mrfusion 11y agoI'd love to see a Wikipedia styled search where people can improve or flag results as they see fit. I wonder if that has been tried. Sure it might not handle the long long tail but the top ten million searches would still be pretty useful.
- deusu 11y agoI think Google once mentioned that each day a surprising amount of searches are unique. They were never done before. If memory serves me correct that number was 30-40%.
- mrfusion 11y agoEven if that's true, 60% is still pretty useful.
- effie 11y agoI'd like to see such search as well. But I think the long tail queries results may not be so bad compared to the current ones, because, especially for obscure queries that end up being searched only a very small number of times per week, the risk of malicious manipulation is (I think) lower and the documents people end up their search with is a good indicator of relevant result.
- jdimov10 11y agoIf it keeps being THAT fast after they've indexed the whole web, I'm switching search providers! :)
- faizshah 11y agoI like it! The explainer tool gives a really cool insight into the results: https://explain.commonsearch.org/ https://explain.commonsearch.org/
- onion2k 11y agoOne thing I hope this project does that Google fails to do is give developers a good API to access search. Google closed down their first web search API and now only give developers access to a limited Custom Search API that's rate limited to 100 queries a day for free with a hard limit of 10k searches - that makes it either very hard to develop anything against or relatively expensive. There are other options (Bing, Faroo, raw access to CommonCrawl) but they're either low quality or hard to work with. A good quality, straightforward, open web search API would be awesome.
- faizshah 11y agoFrom their FAQ: "We will particularly stand out with features that are not in the best interest of commercial search engines: showing less (if no) ads, having an open API, and generally not trying to maximize the amount of time users spend on our service." Seems like right now they are focusing on getting contributors though.
- visarga 11y agoI second this. The search engine could be an API to the web. There are lots of resources that could be pulled out that way.
- mtrn 11y agoI would pay for an API, that gives me access to partial quality crawl content. It would be even better, if the web would be treated, at least for parts, as a digital library and non-profit organizations would recognize the value of access to such a resource and provide it (just as maintenance for roads or public schools).
- sylvinus 11y agoThanks for the feedback! What would be your usecase for that API?
- jclos 11y agoI'm not the original commenter but there could be a big huge case in research. Lots of researchers work on UIs for search, interactive search systems, and query refinement algorithms that are really just abstract layers over an existing search engine. It used to be that we could just overlay stuff over Google, but most search engines nowadays are a pain to work with.
- jasode 11y agoI like the project's goal but as techies, we inevitably want to understand the technical details and how it helps (or handicaps) the search results in comparison with Google. For example, the project's data sources[1] says that the bulk of data comes from The Common Crawl. It looks like the CC is ~150 TB of data[2]. I'm not familiar with google.com internals but various sources estimate that their proprietary crawl dataset is more than a petabyte. (A googler could chime in here with more accurate data.) So it's not as simple as the algorithm for Common Search being "more fair" than the algorithm for Google Inc. The underlying dataset in terms of quantity, recency, rules for the robot, etc all affect the algorithm. This is not a criticism of the project. It is my attempt to understand what is not obvious on the surface level. [1]https://about.commonsearch.org/data-sources https://about.commonsearch.org/data-sources [2]http://commoncrawl.org/2015/12/november-2015-crawl-archive-now-available/ http://commoncrawl.org/2015/12/november-2015-crawl-archive-n... (I'm can't tell if each archive of MM/YYYY is cumulative or an addendum.)
- sylvinus 11y agoHi! Data is indeed as important as the algorithm. Common Crawl is a very good bootstrap but we will certainly need to go beyond once it proves to be the limiting factor. We also hope we can help them improve their data set in the short term by giving them a larger URL seed list.
- libeclipse 11y agoI've tried using different search engines to Google numerous times, but each time I've returned to Google simply because the searches are better. They're more accurate, more relevant, and I very rarely find myself searching more than once to find something. If commonsearch can beat Google in that regard, then count me in. But I doubt it will.
- JohnKacz 11y agoWhen I switched to DuckDuckGo last year I read an interesting comment from someone. The basic idea was that we have all become so accustomed to Google's results and the manner in which we use it (i.e. the way we define our search terms) that it is actually we who must be willing to reprogram our search practices if any competitor is to have a chance to catch up. I'm not sure if I buy that, but I do believe if we don't commit to alternatives it will be next to impossible for alternative search engines to get as good as Google with result quality and relevance. Google simply know too much about me and has performed so many more searches for a rival to outperform them. I still use DDG's 'g!' often but I feel like I'm doing my part to help DDG get better for me and other users.
- djsumdog 11y agoSame here. I say I'd use !g on duckduckgo about 1/3 of the time. If I'm getting really frustrated on a problem and am hacking away, I sometimes just default to the !g. Duckduckgo is now doing localized results (you can choose your region, so it's transparent, unlike Google's). The thing about Google is that, even if you're not signed in, it still tried to present you with personalized results (based on previous searches for that session, your IP, your region ... if you're searching from work; it probably factors that in as well). When people talk about getting to the top of Google results, my response has always been, "Well you need to be more popular and relevant. Also you may be at the top..for some people, but not everyone."
- sylvinus 11y agoI definitely share your usage of "!g", which is why Common Search already supports it ;)
- mynewtb 11y agoSeeing how the founder is the same who founded Jamendo which later was turned into a sad, user-unfriendly attempt to make money with freely licenses music (destroying its community in the process), how can I trust commonsearch not to be a waste of time and attention?
- sylvinus 11y agoSo much anger! I left Jamendo 6 years ago so I have limited influence on what they do now, unfortunately. Common Search is a nonprofit and 100% open source so it is fundamentally different.
- vmorgulis 11y agoThe demo if pretty fast.
- Fastidious 11y agoI do not sense any anger on the OP comment. It sounds like a legit, real concern.
- sylvinus 11y agoWell I actually share some of that anger so I'm sorry if I read too much into "destroy" and "sad". Common Search is forkable by design so it should hopefully stay on course one way or the other!
- deleted 11y ago[deleted]
- barryhunter 11y agoIt's up to you if want to take the risk :) There is always a risk in trying something new. It might never pan out, or disappear in a puff of smoke. Nothing would happen if nobody took a little risk.
- dmvaldman 11y ago
- rmc 11y agoI'm trying to find out from their website, but it's unclear. Are the servers hosted in the USA? And will the organisation be incorporated in the USA? If you're talking about privacy and transparency, it's better to operate in a place bound the European Charter of Fundamental Rights, rather than the US Constitution, because the former gives people much more rights with their data, how it's used, etc.
- sylvinus 11y agoIt is very early so we are not yet incorporated. The issues you mention will definitely be taken into account! At scale, we'd probably have multiple legal entities in different countries anyway, like Wikimedia.
- rmc 11y agoThanks for your concern. However I was under the impression that Wikimedia was a US organisation, with some local chapters. Perhaps incorporate in a EU country from the start?
- betolink 11y agoTotally on point, if privacy is a concern the company should not be inc. on American soil. I'm also having a hard time trying to find the roadmap for infrastructure, how many machines are running or will be expected to run. Scaling is not trivial and I don't see how that would be addressed.
- sylvinus 11y agoRight, there is no specific roadmap for infrastructure yet or view of what servers are currently running. I will add those things on our Operations page, thanks! https://about.commonsearch.org/developer/operations https://about.commonsearch.org/developer/operations
- struct 11y agoNeat, I was working on a project to give a full programmatic keyword index to the contents of the common crawl, but I guess there's no need! It's very exciting to consider what kind of applications you can build with this.
- ocdtrekkie 11y agoThis sounds awesome. Speaking of building AIs/bots and such in your FAQ, the lack of a good open API for search is probably what gates that market to Google and Microsoft and such... That nobody else can just tap a search engine. I'd love to be able to connect to this for queries at some point.
- tonylxc 11y agoI'm particularly interested in the discuss forum. Is it an open source one or built yourselves? Thanks!
- sylvinus 11y agoOh no, one project at a time :) We used http://www.discourse.org/ http://www.discourse.org/, which I recommend.
- PaulHoule 11y ago"nonprofit" for me is a bad smell. I.e. the problem of sustainability, which for nonprofits is all about the money and not about carbon or solar energy, rainbows, plutonium or any of that.
- whazor 11y agoI think people might underestimate the power of an open source search engine. In my eyes it is like wikipedia versus the old paper encyclopedia books. Improvements to search results in Google are done by a relatively small amount of people from Google. Google decides where you buy, what you think and how you live. Behind their algorithms they probably have made dozens of subjective choices. Public debate, more attention to details, and open politics are as I see it, great tools to improve search engine quality.
- schoen 11y agoAnd even if organizations smaller and less sophisticated than Google can't match the kind of search quality that Google has achieved, the autonomy and transparency issues are still important. An interesting metaphor for search engines' power is in http://james.grimmelmann.net/files/Library.markdown http://james.grimmelmann.net/files/Library.markdown
- praxulus 11y ago>Improvements to search results in Google are done by a relatively small amount of people from Google How many open source projects log more engineering hours than Google's search team? It's the flagship product of one of the largest corporations in the world.
- whazor 11y agoWikipedia has 133,621 active registered users, and 27,755,916 users. Furthermore, Wikipedia has 819,043,068 page edits. While Google will probably have better and more engineers, but they rely on usage statistics, and not experts in the specific search domains.