Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dor_jack_2
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
dor_jack_2
9y ago
Loop/spam prevention was done by mixnode, I'm not sure how they do it. The data does not follow a DFS or BFS pattern so pages/site varies greatly by a host's server capacity and anti-crawling configs. There was a minimum
2.
▲
by
dor_jack_2
9y ago
For our purposes Common Crawl's corpus was missing too many websites (possibly due to robots.txt configs of websites) Also we needed some deep coverage which CC could not provide.
3.
▲
by
dor_jack_2
9y ago
We tried a bunch of technologies like Nutch, Heritrix, Storm Crawler, ... eventually settled on Mixnode and since it's a 'cloud platform' we didn't really have to change anything. As for processing the data we crawled, w