3 ms·
If anyone's interested in web crawling technology, check out Heretrix [1], been around since 2004 and while not the most performant it has incorporated many res
by kingforaday 3y ago
If anyone's interested in web crawling technology, check out Heretrix [1], been around since 2004 and while not the most performant it has incorporated many responsible disciplines in the design and as this article pointed out, WARC format.
1. https://heritrix.readthedocs.io https://heritrix.readthedocs.io
- fosstrack 3y agoSecond that. Anyone interested in studying web crawler tech should definitely take a look at Heritrix. I had used it extensively when it was still in 2.x. They got so many things right about writing well-behaved and fault tolerant crawlers. Plus the code is very modular, and extensible, if you know some Java. The other popular option then was Apache Nutch, but it had too much hadoop baggage.
- marginalia_nu 3y agoHadoop is a bit of a nuisance in this general corner of Java. It's got a propensity for integrating deeply with cluster adjacent technology in a way that is very difficult to root out. Kind of a pity since it has the effect of making things that could be very easy, such as reading and writing parquet files, much harder than it needs to be in Java.
- abracadaniel 3y agoIts performance shines in larger scales. It’s designed for politeness to individual domains, but scales out well for very wide crawls of many domains. It’s pretty much endlessly configurable, but not the easiest to learn.