14 ms·
Show HN: Open-source search engine with 2bn-page index
- deleted 10y ago[deleted]
- outpan 10y agoAwesome job! For the life of me I can't figure out how you manage to crawl over a billion web pages (even in 2-3 months), index the data and run the server with €300 per month. Especially the crawler part...
- vcool07 10y agoAny specific reason you've used pascal ? I thought that language got extinct long ago.
- deusu 10y agoIt's alive and well. The TIOBE index still lists it ahead of Ruby, Swift, Objective-C, GoLang... And I started this software 20 years ago. Granted, a LOT of the software has changed since then. But I don't see a reason to throw away existing code unless it is in need of so much change that rewriting from scratch would be easier. And even then I might stick to what I know best, and what fits best with other parts of the software.
- supersan 10y agoHi, I find the Blog more interesting right now since I hope to find write-ups about how you were able to manage such a herculean task on your own? Crawling 2bn pages could take forever and could generate a huge bandwidth bills, so any lessons you learnt, pitfalls you faced, etc would be a great read.
- deusu 10y agoSome issues that appeared over the years: Block outgoing connects to local IP nets in your firewall. Otherwise your hosting provider might think you are trying to hack them. Apparently there are a lot of links out there that point to hosts which resolve to private IP ranges. Another problem with following links is that you are bound to run across some that are malware command & control servers. Had several complaints to my ISP after authorities took over control of one and used the C&C server's domain as a honeypot. My crawler is on a whitelist now. I had one person who vehemently complained that I was trying to hack him, because the software downloaded his robots.txt. I'm NOT kidding! :) Make sure your robots.txt parsing is working correctly. I had an undiscovered bug in the software at some time which basically caused it to think everything is allowed. Luckily someone was nice enough to let me know. And he was really nice about it. And he would have had every right to be angry. A major bottleneck is DNS queries. Run your own DNS server and even cache the hostname/IP pairs yourself. Do not even think about using your IPS's DNS server. If you bombard them with 100+ DNS requests/s then they WILL be angry. :)
- webtechgal 10y ago> Run your own DNS server and even cache the hostname/IP pairs yourself. This[1] might be a useful resource to get started: [1] https://scans.io/ https://scans.io/ (Register and download the IPv4 Address Space data file to use as an initial cache and then append/update as you go.)
- deusu 10y agoBookmarked. Thanks!
- rbjorklin 10y agoWhat makes this better than https://duckduckgo.com https://duckduckgo.com ?
- diggan 10y agoNot saying that it's better but one of the main selling-points of DeuSu seems to be that it's fully open source and independent search index. Duckduckgo, if I remember correctly, is not 100% open source and get their search index from Yahoo (or maybe Bing, not sure)
- kowdermeister 10y agoIf it's not good, the it doesn't matter if it's OSS or not.
- deleted 10y ago[deleted]
- anewhnaccount 10y agoGood for what? Even though this isn't good for use as an every day general purpose search engine, it could be good for a particular use case perhaps with some adaptation or for learning from.
- kowdermeister 10y agoI don't know why would people use it to be frank. Lot better alternatives exists. > it could be good for a particular use case Namely? > or for learning from. The author admitted in the github readme that the code quality is rather bad. I also don't see a link to the search index, the only valuable component of this project.
- deusu 10y agoI will publish the index for download in a few weeks. I'm currently working on the documentation. Oh, and I will publish the raw crawl-data too. Everything together is about 2.5tb. There is also a free API in beta-test right now. Will probably be ready for official release next week.
- fnord123 10y agoIt's written in pascal. Neat. However, it's not very good. If I search for "banana" I get information about a sex shop rather than about bananas.
- 0xmohit 10y agoAppears to be a "smart" search engine. Tries to infer a lot (maybe based on data collected earlier).
- kowdermeister 10y agoStrange, Wikipedia article is not on the first page and don't blame me for searching something non German thing :) https://deusu.org/query?q=berlin https://deusu.org/query?q=berlin
- semi-extrinsic 10y agoIt's pretty obvious that Google et al. do a lot of "custom" filtering like prioritising Wikipedia, removing porn from "obviously non-porn" searches etc. (That "Berlin" search gives porn as the 8th result.)
- gkst 10y agoI doubt that Google prioritizes Wikipedia deliberately. Wikipedia has tons of backlinks, authority, trust, typically a high text to html ratio, probably a low bounce rate. Moreover, it is fast, works well on mobile and on and on. It's is just a very well done and useful site for users and search engines.
- kowdermeister 10y agoMy thoughts as well. They don't need special treatment to be in the top 3.
- allendoerfer 10y agoWikipedia being ranked high is even an indicator for SEOs that an affiliate niche is not very competitive.
- throwaway13337 10y agoAlternative general purpose search engines are an exciting idea. It seems a lot like we're about the time when yahoo was dominant and searching was sort of awful. When you searched, what ranked highest was market-driven sorts of stuff. Right now, for topics normal people search for - not techies -all you get are content farm sites with js-popups asking for your email address. Try searching for anything health related, for example. We've regressed. My own half-solution is to look only for sites that are discussions - reddit, hn, etc. It could be better. A search engine that favored non-marketing content could really steal some thunder. This doesn't look like that, but maybe its a start?
- samstave 10y agoWhy not have a search engine with "sub-reddits" that can be subscribed to... Whereby - a site would self-identify as being in a particular genre, say "healthcare" - and I could launch a tab to the engine and set my sub to "health, health-tech, healthcare, medicine, etc.." and then do my search and only those sites that set their category will show up in that search - but if I dont find my search, I can then easily slide out to other areas where I may not have thought what I was looking for would have identified with. Further - any post by any company/site could individually been given a topic to self-declare as... thus even if the company or site isnt necessarily in that space - their page or object could at least be a part of that result ranking.... Or has this been tried/found to be stupid?
- allendoerfer 10y ago> Or has this been tried/found to be stupid? You are describing the keywords meta tag. While it is often told that competitors before Google did not use something like PageRank, which is not true, Google's PageRank algorithm was better and cheaper than the competitors' and effectively killed your idea 20 years ago.
- samstave 10y agoAppreciate the insight.... But I find it slightly ironic that people are bitching about PageRank having slightly some issues with respect to the specificity of what they are searching for... meaning that even though "killed this idea twenty years ago" we are coming back to the same problem... Is that perhaps just due to the volume of info that is available on the web and the much more complex way we have categorized (mentally, not digitally) all the knowledge and information thats out there now?
- laurent123456 10y agoThey need to filter porn out of their search results (even for common queries like "hat", there's only porn) and perhaps be more resilient to SEO techniques since it looks like there's lot of spam on top results. Queries with common words such as "cat" return almost only irrelevant results. I'd really like to see that kind of project working as a good alternative to Google, but as it is it's not really usable.
- deusu 10y agoI hadn't even thought about that. But it should be pretty easy to do in post-processing. I just have to take a list of "porn" keywords. If none of them occurs in the query, but in a search-result, then that result gets downranked.
- laurent123456 10y agoYes I guess filtering them out would at least make the website SFW, and it would make it easier to show it to people. The issue seems to happen mainly with common words (which results also appear to be polluted with heavily SEO-ed websites). I've also searched for less generic things like "xperia z5" and the results looked good.
- deusu 10y agoI have the filter implemented now. It's not perfect yet, but it already filters out a lot of the NSFW stuff. Unless you explicitly search for it. I'm gonna further improve this over the next days. Right now it's just a quick'n dirty hack. :)
- NKCSS 10y agoFun, but overal quality seems a bit lacking. When I search myself; the top 10 results don't even have my last name ('Kusters') and just shows pages that have the word 'Nick'. I suppose you don't use a form of LSA to score the search results? Maybe it's too specific, but afaik mainstream search engines seem to give somewhat consistent results here. https://deusu.org/query?q=nick+kusters https://deusu.org/query?q=nick+kusters Looking at the code (https://github.com/MichaelSchoebel/DeuSu/ https://github.com/MichaelSchoebel/DeuSu/) I notice that you have ranking modifiers based on the .tld; why not store the reported content language and score based on that? Isn't that more relevant?
- deusu 10y agoIn my experience this is usually caused by the fact that even 2bn pages aren't that many nowadays. The index needs to get bigger to better find (and rank) long-tail results like queries like this.
- scandox 10y agoEvery time I see new search engine projects I remember this: https://en.wikipedia.org/wiki/Cuil https://en.wikipedia.org/wiki/Cuil I note that Dr Anna Patterson is back with Google. She wrote this in 2004: http://queue.acm.org/detail.cfm?id=988407 http://queue.acm.org/detail.cfm?id=988407
- hvo 10y agoI am not sure many of the issues Dr. Anna Patterson raised here are applicable now.Web is way different now compare to 2016.
- mstolpm 10y agoIn addition to the lack of removing porn and the ordering of the results not priorizing "quality" sources, some of the indexed site data is at least 4-6 months old and has heavily changed since the last crawl. I even got 404 errors. That makes it very hard to really find use in the project other than for academic interest.
- deusu 10y agoA fresh recrawl is currently running. Should take about 2-3 months. Newly crawled data will gradually replace older data during that time.
- webtechgal 10y agoGreat work, congrats. :-) Here is some input based on my experience building a similar project at my former company. (We did not quite get to 2B pages, but were close to ~300M): For creating a really viable (alternative) search engine, the freshness of your index is going to be a fairly important factor. Now, obviously, re-crawling a massive index frequently/regularly is going to need/consume some huge amounts of bandwidth + CPU cycles. Here is how we had optimized the resource utilization: Corresponding to each indexed URL, store a 'Last Crawled' time-stamp. Corresponding to each indexed URL, also store a sort-of 'crawl-history' (If space is a constraint, don't store each version of the URL, store only the latest one). On each re-crawl, store two data fields: time-stamp and a boolean if the URL content has changed since last crawl. As more re-crawl cycles run, you will be able to calculate/predict the 'update frequency' of each URL. Then, prioritize the re-crawls based on the update frequency score (i.e. re-crawl those with higher scores more frequently and the others less frequently). If you need any more help/input, let me know and I'll be happy to do what I can. HTH and all the best moving forward.
- webtechgal 10y agoWe had also (obviously) built a (proprietary) ranking algo that took into account some 60+ individual factors. If it can be of any help, I'll create a list and send it to you.
- malinens 10y agoworks really fast!
- deusu 10y agoThx. But all the traffic from here is currently driving the servers to their limit. Queries are already slowing down a bit because of imminent overload. Usually the average query takes about 250ms. Currently the average is at 334ms.
- swiley 10y agoThe site's interface is just incredibly pleasant compared to Google.com. I really hope the author sticks with it. Unfortunately I'm not sure it's usable right now, searching "group theory Wikipedia" never brings up a Wikipedia page (although maybe I should just be directly searching Wikipedia if that's what I wanted).
- DanBC 10y agoDuckDuckGo's approach of !bang searches, making duckduckgo the place[1] I go when I want to search another site, is really useful. [1] It's my default search engine in Chrome, so I use bang searching in the address bar.
- Cyph0n 10y agoSame here. The problem is that I find myself using `!g` way too often... I guess I'm not used to the DDG results page.
- amirouche 10y agoddg is my primary search engine, it takes time but you get use to it. If what you are looking for is mostly on HN, SO or wikipedia it works quite well.
- micwo 10y agoDeusu can't find deusu (or deusu.org) https://deusu.org/query?q=deusu https://deusu.org/query?q=deusu
- deusu 10y agoAnd why should it? You are already at the destination. No need to find it. :)
- micwo 10y agoTry to find any other site by url: https://deusu.org/query?q=news.ycombinator.com https://deusu.org/query?q=news.ycombinator.com
- 0xmohit 10y agoTry to find `2 + 2 = 4`: https://deusu.org/query?q=2+%2B+2+%3D+4 https://deusu.org/query?q=2+%2B+2+%3D+4 Even https://deusu.org/query?q=2+%2B+2+%3D+5 https://deusu.org/query?q=2+%2B+2+%3D+5 didn't yield any results. I was under the impression that it'd a message: 2 + 2 = 5 for very large values of 2.
- 0xmohit 10y agoEarlier discussion: https://news.ycombinator.com/item?id=9122397 https://news.ycombinator.com/item?id=9122397
- gkst 10y agoPascal is an interesting language choice. I think it is the 1st time I see an open source project that is actually used in production written in Pascal.
- ashitlerferad 10y agoAnother open source search engine: http://yacy.net/ http://yacy.net/
- ytjohn 10y agoThanks, I was trying to remember that one. I think that for any new, non-profit, search engine to be viable, it has to be decentralized. deusu.com takes 2-3 months to crawl 2bn pages. Yacy claims to be at 1.4bn. I don't know how long it takes for that index to get refreshed, but it has 600 peer operators. Even if Yacy has a weaker indexing algorithm, I imagine that 600 peers, each crawling and contributing their own set of sites must be faster than a single deusu node. Yacy is also quite a bit more resilient. I will say that I don't buy Yacy's "no censoring" statement. If I was a bad actor, I could run yacy on a computer with false dns and false certificates, and yacy could index my fake content with official looking URLs.
- yati 10y agoLooking at the source code took me back to days when I used to do stuff in Delphi :) Neat project -- Loads of room for improvement, but a great initiative!
- ommunist 10y agoDeuSu does not crawl social pages it seems. No traces of linkedin profiles and no facebook. From a certain point of view - this is a good thing.
- pmontra 10y agoWritten in Delphi. I might be wrong but I don't see many people downloading and working on it. 30 day free trial and then you have to pay for the development environment. IMHO it's a non starter for an open source project but if it's the only language the author is comfortable with, well that's OK.
- deusu 10y agoOriginally it was written in Delphi. But I now use FreePascal for the development. I'm even compiling both Windows and Linux versions on my Linux machine.
- pmontra 10y agoGreat choice! Thanks.
- RobAley 10y agoIt appears to now have moved over to FreePascal, which is the free Open Source delphi look-a-like.
- ommunist 10y agoDeuSu seems not indexing Cyrillic part of the Internet, and cannot give you insights for Greek, try https://deusu.org/query?q=ελιά https://deusu.org/query?q=ελιά . Is it Latin ANSI only index?
- deusu 10y agoOnly ASCII and German umlauts (äöüß) at the moment. The parser needs rewriting. It was originally written in pre-unicode times. :)
- tychuz 10y agoAnd all javascript related questions still have w3schools as first result, god dammit.
- gkilmain 10y agoI think for newbs who want to learn the fundamentals of web dev w3schools is a good resource. Even the people over at w3fools admit it. For a deeper dive though clearly MDN is the winner.
- skykooler 10y agoIt shows snippets of the web pages under each result; however, generally not the particular snippets that contain the search term. I would think that would be useful.
- deusu 10y agoYes, it would be better. The snippets are currently the first 255 characters of the page's text. For snippets to be customized to the search term, I would have to store all the text of the page. And that would require a lot more disk space. Space that I can't afford at the moment.
- vain 10y agoGoogle's secret ingredient to stay relevant and informational is Wikipedia. Deusu on the other hand seems to weight words in urls highly. If you search for scientology only on Deusu, you might end up wearing a funky hat https://deusu.org/query?q=scientology https://deusu.org/query?q=scientology
- Taek 10y agoIs the two billion page index open source? I've been thinking a lot about days recently. Seems to me like Pandora's box is open. Google knows where you live, where you eat, what your fetishes are, all of your sexual partners. Facebook knows most of those things to, via different methods. And if you run Windows Microsoft probably has access to most of that as well. Apple will too, because if they don't they won't be able to compete. Tesla, Uber, Waze also have a huge amount of data on your life. Everyone is pushing the envelope on how much data they are collecting, and the companies which collect more data will compete better. As tech gets better we will increasingly be unable to resist sharing our whole lives with the companies who are powering modern living. Even worse, there's a huge monopolization effect to having data. Nobody else has anywhere near as much data as Google. That means nobody else can compete. Nevermind the engineering, your algorithms can be 2x as good but you won't have 0.1% the data as a company with billions of daily users. So Google and Facebook are left untouchable. Microsoft, Apple, and maybe Amazon can get in range. Is there anyone else? We can fight back by giving up the privacy war and blowing the doors open instead. Take your data (as much as you dare) and make it public. Let every startup have access to it. Let every doctor have access to it. Give the small players a fighting chance. That does mean a massive cultural shift. It means your neighbors will be able to look up your salary, your fetishes, your personal affairs. It's a big deal. I don't see any other way out of this though. Surveillance technology is getting better faster than privacy technology, because surveillance tech has the entire tech industry behind it. Smarter phones, smarter TVs, smarter grocery stores, smarter credit cards, smarter shoes... smarter everything. Privacy is melting away and we aren't getting it back.
- deusu 10y agoThe software is already open-source. A free search API will be fully available probably next week. It's in testing already. It's just a matter of putting the finishing touches on the documentation. And the crawl- and index-data will be available for download in a few weeks. It's also just a matter of documenting the data-format. BTW: I disagree with your points about privacy. I see DeuSu as a way of fighting back.
- deleted 10y ago[deleted]
- jbb555 10y agoI think projects like this are really important because they help reduce the impression that big server projects are only meant to be done by big companies. The internet is becoming a content consumption medium for many people. I'm not sure I'll use this, but I'll try to... it all depends on how good it is. But I approve of the project so I sent a (very) small bitcoin donation to hopefully help fund it for a few more minutes :)
- deusu 10y agoThank you! Depending on who you are (there were 2 bitcoin donations today), you funded either about 18 or 28 hours of operations. :)
- CM30 10y agoWell, I admire the work behind it, and I think the idea is good (especially how having this open source means multiple sites can build on the same data set and get it more and more accurate over time). But I have to be honest and say that it's just not working for me. I type in Reddit, and it shows links to the NSFW subreddits instead of the main site or anything else on it. Typing in Wikipedia gives me the Dutch version of Wikipedia. Mario Wiki? The page on Mario Wiki about Mario Clash, then the Smash Bros Wiki and a bunch of SEO spam pages. Pokemon Go gets me no relevant results at all. Certainly not anything official, that's for sure. It's a decent start, and having 2 billion pages indexed is pretty impressive for a project like this as it is, but it's just not really usable as a search engine just yet.
- ccleve 10y agoYou get really good performance on not much hardware. Can you share some technical details? - file formats, particularly the postings - query evaluation strategy - update strategy I poked around in the source code a bit, but couldn't find these things.
- deusu 10y agoFile formats will be documented when I publish the data-files in a few weeks. What do you mean with postings? The main index is split into 32 shards (there is also an additional news-index which is updated about every 5-10 minutes). Each shard is updated and queried seperately. The query actually runs 2/3 on a Windows server and 1/3 on a Linux server. The latter in Docker containers. I want to move everything to Linux over time. Query has two phases. First only a rough - but fast - ranking is done. Then the top results of all shards are combined and completely re-ranked. This is basically a meta search engine hidden within. First query phase is in src/searchservernew.dpr, and the second phase is in src/cgi/PostProcess.pas.
- ccleve 10y agoThank you. "Postings" is another word for the format of the doc ids and related information in the inverted file. A google for "inverted index postings" will turn up a bunch of references.
- billconan 10y agoI searched "meta programming c++" and the top returns are all about java. I'm curious, is it expensive to run a search site like this?
- deusu 10y agoCurrently €300/month. More details on https://deusu.org/donate.html https://deusu.org/donate.html
- amirouche 10y agoDid you think about database dump of popular services like HN, SO or Wikipedia to speed up crawling and revelance?
- deusu 10y agoYes. I have downloaded several data dumps, but haven't gotten around to import them yet.
- rshm 10y agoAs of aug 16, common crawl has 1.73n pages. For the complimentary set of urls, if any benefit you can use their data dump as seed. If the metadata (such as last modified) size of your index is small enough to upload to aws, you can also reduce your re-crawl efforts when they have a fresh release.
- greglindahl 10y agoIt doesn't have to be small to donate to Common Crawl, they have a free S3 bucket.