4 ms·
Making the software scale to N servers is one of the reasons behind the planned rewrite. I'm already considering making "special" indexes for blogs, news etc.
by deusu 12y ago
Making the software scale to N servers is one of the reasons behind the planned rewrite.
I'm already considering making "special" indexes for blogs, news etc. like you mentioned. In fact that same suggestion came up in a conversation I had with someone today.
The crawler/indexer is pretty fast as it is. I could crawl about 2 billion pages/month for a cost of about 800€ / 900US$ per month. That's on a 3.4GHz quad-core machine with a 1gbit/s connection and 32gb RAM. I tested that once. Works fine.
I can only imagine that your parser does a lot more than mine does. Crawling is CPU-bound for me. So the more processing the parser has to do, the slower the crawling would be.
- ChuckMcM 12y agoExcellent! The crawler does get CPU bound so having multiple machines (and/or cores) makes it go faster. We (as in Blekko) also maintain a list which we call the "crawl frontier" which are pages we know about but haven't seen enough relevant data about to decide whether or not they are "worth" crawling. Although some new stuff we just built will help with that. Much of what we crawl is shared with the common crawl folks so that can help you with your choosing perhaps. Also that corpus provides a good way to test indexing and ranking algorithms. Things I like to test are queries like "Who are the Cardinals?" which can catch both sports references (Arizona and St Louis Cardinals) and church references (Pope Francis just named/blessed/appointed a bunch of new Cardinals). Often times news sources or blogs can help identify which interpretation of an ambiguous word is probably more relevant at the moment.
- deusu 12y agoI usually start a crawl with a list of the top one million websites according to Alexa.com. Unfortunately that list doesn't seem to be available for download anymore, so I have to make due with the latest one I have which is from April 2014. After that URLs get crawled basically in a first-come first-serve order. I simply crawl them in the order I find them. Interpreting ambiguous words for ranking is beyond what I can do at the moment.
- frik 12y ago> URLs get crawled basically in a first-come first-serve order I saw some references to round-robin in your code, I suppose you use it to hit a domain only every x milliseconds during a crawl. Maybe additionally use Open Directory Project as start list: http://dmoz.org http://dmoz.org