3 ms·
Check out Nutch, not sure if its exactly what you want though. It's in Java not Python, but it works with Hadoop quite nicely
by aonic 19y ago
Check out Nutch, not sure if its exactly what you want though. It's in Java not Python, but it works with Hadoop quite nicely
- dshah 19y agoI'm going to vote for Nutch too. Have heard good things. Also, if someone on this list wants to work on a cool web spidering project (probably using Nutch), send me a message. I'm looking for someone.
- surya 19y agoI'm interested...
- groovyone 19y agoThanks. Still would like to use Python to be honest (any python suggestions?), but I'll give this a go. Going to do some more research and might post back findings if anyone would be interested in critiquing them. I'm creating this startup from scratch so if there is anyone interested in the crawler side of things I'd be happy to chat either about collaboration or sharing ideas.
- konsl 19y agoIf you're looking at building your own crawler in Python from scratch, here's a benchmark of SGML parsers: http://72.14.205.104/search?q=cache:LYoRD1GTP2UJ:www.oluyede.org/blog/2007/08/25/sgml-python-parsers-benchmark/+benchmark+python+parsers&hl=en&ct=clnk&cd=1&gl=ca http://72.14.205.104/search?q=cache:LYoRD1GTP2UJ:www.oluyede... We've been playing with sgmlop (http://effbot.org/zone/sgmlop-index.htm http://effbot.org/zone/sgmlop-index.htm) for parsing and urllib2 (http://docs.python.org/lib/module-urllib2.html http://docs.python.org/lib/module-urllib2.html) for fetching.
- sarosh 19y agoI concur with the Nutch vote; but more specifically, take a look at the crawler code written in the src trunk for use with Hadoop. That is probably a good place to start. Also worth a look is Heritrix (crawler for archive.org). http://sourceforge.net/projects/archive-crawler http://sourceforge.net/projects/archive-crawler Sadly, this too is written in Java. The only Python one I am aware of for which code is available is: http://sourceforge.net/projects/ruya/ http://sourceforge.net/projects/ruya/ Edit: You might also want to take a look at http://wiki.apache.org/hadoop/AmazonEC2 http://wiki.apache.org/hadoop/AmazonEC2 Edit2: Polybot is another Python based crawler, but no code. However, the paper has some interesting ideas: Design and Implementation of a High-Performance Distributed Web Crawler. V. Shkapenyuk and T. Suel. IEEE International Conference on Data Engineering, February 2002. http://cis.poly.edu/westlab/polybot/ http://cis.poly.edu/westlab/polybot/
- inovica 19y agoGood response. We've created a basic crawler in Python, but are looking for something more powerful too. Heritrix above looks good