5 ms·
Please keep in mind that this is running on one server. Google has what? Several hundred thousand servers? Yes, it needs to get better. But it's a good start,
by deusu 12y ago
Please keep in mind that this is running on one server. Google has what? Several hundred thousand servers?
Yes, it needs to get better. But it's a good start, I think.
- fiatjaf 12y agoIt is not a good start. It needs to get much better to be compared to Google, and it will not, if it is a one-guy work and you haven't discovered a new fantabulous almost magical new algorithm for crawling and indexing. Keep in mind that Google is not a propaganda machine, it really really _really_ solves problems and finds things. Any search engine willing to compete will have to do this. I'm not saying it is impossible, but that it is not sufficient to be "advertising-free" to be good.
- rileyteige 12y agoIn no way do I intend to discourage your efforts. I was offering an example in which your search engine seems unable to identify the context that I had in mind. What are some of your plans for search engine improvement? Ad-free aside, why should I fund your project? DuckDuckGo seems to offer itself up as a great alternative to Google and I have enjoyed its clean UI immensely for several months now. What is my incentive to help drive your project?
- deusu 12y agoIt's open-source: https://github.com/MichaelSchoebel/OpenAcoon https://github.com/MichaelSchoebel/OpenAcoon This is currently written in Pascal. I'm in the process of rewriting it JavaScript/Node.js. First part for the rewrite is the ranking. That should be done in about a week or two.
- NhanH 12y agoPascal! Now that's a name I haven't heard in a while. You should write a blog post about how you write a search engine in Pascal. That would at least reach top spot on HN!
- deusu 12y agoGiven that I'm in the process of rewriting it in JavaScript I'm not sure about that... On the other hand it's probably worth a try. Thanks for the suggestion.
- frik 12y agoHe wrote a blog post there: http://geekregator.com/2015-01-05-welcome_to_geekregator_com.html http://geekregator.com/2015-01-05-welcome_to_geekregator_com...
- xai3luGi 12y agoHave you considered joining the YaCy community? It is another open source search engine. http://www.yacy.net/ http://www.yacy.net/
- deusu 12y agoI know of YaCy. It's written in Java which I'm almost completely unfamiliar with. But I know and follow the YaCy-community. Like myself the lead person of YaCy is located in Germany. We even share the same first name. :)
- fiatjaf 12y agoWhat about Faroo? http://www.faroo.com/ http://www.faroo.com/
- xai3luGi 12y agoInteresting but there is approximately zero technical information about how it works on their website. Minimal code release. Not a distributed network of independently run nodes (like Yacy or Tor).
- zz1 12y agohttp://www.faroo.com/hp/p2p/p2p.html http://www.faroo.com/hp/p2p/p2p.html ? Looks like some ideas are the same featured in https://www.blippex.org/ https://www.blippex.org/
- e12e 12y agoHaving a look at the source ... why are you rewriting this in node/js? Wouldn't Ada (or nim) make more sense, given the initial design is implemented in Pascal?
- deusu 12y agoMany of the design choices that I made many, many years ago aren't valid anymore. So it won't just be a port to a different language, but a complete redesign and rewrite. Which means that regarding the amount of work, it doesn't really matter which language I use. JavaScript seems to have the biggest community at the moment. Plus I kinda like Node.js. That's why I'm going that way.
- ChuckMcM 12y agoIts a fun way to look at what might be involved in writing a search engine and to start experiencing the myriad of ways in which people "search" and what they expect when they do. I encourage you to continue in your efforts, and I will add some suggestions on how you might proceed which will get you further along. First, in order to scale, it will have to be able to run on 'n' servers. So look at ways of breaking up your index across multiple machines and operating on the index in parallel. Second, crawling can often be a bigger challenge than providing search results, so work on building a system that can crawl a bunch of different web sites. Folks like Octopart have shown that topical search engines are really useful. So consider putting together a billion document index of say "news sources". In the 'everything that is old is new again' theme, consider building simply a 'blog search engine' which is one that focuses on various blog content. The old Technorati did that and later was subsumed and its become harder for blog writers to rank in Google's results, so perhaps there is an opening for a place to search about who is talking about 'x' for some X. I happen to know you can run a pretty decent crawler and indexer for about $2M a year :-) so consider targeting your donation rate to hit about $150,000 a month.
- deusu 12y agoMaking the software scale to N servers is one of the reasons behind the planned rewrite. I'm already considering making "special" indexes for blogs, news etc. like you mentioned. In fact that same suggestion came up in a conversation I had with someone today. The crawler/indexer is pretty fast as it is. I could crawl about 2 billion pages/month for a cost of about 800€ / 900US$ per month. That's on a 3.4GHz quad-core machine with a 1gbit/s connection and 32gb RAM. I tested that once. Works fine. I can only imagine that your parser does a lot more than mine does. Crawling is CPU-bound for me. So the more processing the parser has to do, the slower the crawling would be.
- ChuckMcM 12y agoExcellent! The crawler does get CPU bound so having multiple machines (and/or cores) makes it go faster. We (as in Blekko) also maintain a list which we call the "crawl frontier" which are pages we know about but haven't seen enough relevant data about to decide whether or not they are "worth" crawling. Although some new stuff we just built will help with that. Much of what we crawl is shared with the common crawl folks so that can help you with your choosing perhaps. Also that corpus provides a good way to test indexing and ranking algorithms. Things I like to test are queries like "Who are the Cardinals?" which can catch both sports references (Arizona and St Louis Cardinals) and church references (Pope Francis just named/blessed/appointed a bunch of new Cardinals). Often times news sources or blogs can help identify which interpretation of an ambiguous word is probably more relevant at the moment.