4 ms·
Ask HN: Is there a great open source crawler?
I want to crawl a site, like foxnews.com and find all the URLs that match a pattern.
A pattern like:
http://www.foxnews.com/\w+/\d+/\d+/\d+/.*/<p>That would find all URLs like:
http://www.foxnews.com/world/2010/06/02/report-natalee-holloway-suspect-sought-murder-peru/
I know I could do this myself. I also know it's a seemingly easy problem, that is actually quite hairy.
I'm hoping there's an open source crawler that I can point to a start page and say "Find all the URLs that match this pattern.".
I know there are dozens of crawlers out there. I just don't know if there's one or two that are really great. I'm really hoping there's a modern/simple/fast one that would be good for this purpose.
Does such a thing exist? If not, are there any great documents detailing the common problems and how to solve them?
Thank you.
- jdrock 16y agoTooting my own horn, but this takes all of 1 minute in 80legs if you've got the right regex.
- haihai 16y agoI considered 80legs. The problem is that I don't want to do this for just one site. I may have dozens. I can run this off my own servers/bandwidth (which I already pay for) for much cheaper than using 80legs. If 80legs was purely usage based and cheaper I probably would have used it.
- jdrock 16y agoUnderstood. We do work with folks on custom per-use plans, but we just need to figure out what your requirements are so that we can see if there's some way for us to help.
- yourabi 16y agoNone of the ones I know are that simple - but Check out Heritrix at http://crawler.archive.org http://crawler.archive.org and Nutch at http://nutch.apache.org http://nutch.apache.org Also worth checking out is http://80legs.com http://80legs.com
- Ledio 16y agoNutch has a full on web crawler, a lot of features, and it scales pretty well. You can white list or black list URLs as you see fit, and filter out unwanted content.
- ericwaller 16y agoScrapy (http://scrapy.org/ http://scrapy.org/) matches your description pretty well. You can specify which urls to crawl with regular expressions and then provide a bit of code to do some data extraction.
- iworkforthem 16y agoThere a quite a few open source web crawlers, in Java, PHP, etc.. Based on what you mentioned, Nutch is alright... But not sure how are you going to get all those unstructured information out and make it structured.