Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ahadrana
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
ahadrana
15y ago
Hi, you can view our terms of use at http://www.commoncrawl.org/about/terms-of-use/full-terms-of-... . We adhere to the robots.txt standard, try to do all our crawling above board, and (strictly personal opinion here) we are definitely not
2.
▲
by
ahadrana
15y ago
We hear you. Could you define some criteria as to the type and size of sample data you would like to see? We are working on producing more targeted/limited collections, like perhaps all most recently published blog posts etc.
3.
▲
by
ahadrana
15y ago
The pagerank and other metadata we compute is not part of the S3 corpus, but we do collect this information and probably will make it available in a separate S3 bucket in Hadoop SequenceFiles format. Be aware that our pagerank will probabl
4.
▲
by
ahadrana
15y ago
Hi I work at commoncrawl. We have spent our time (in 2011) improving our algorithms, and hopefully this effort will start to show real results (with respect to crawl frequency and relevancy) in 2012. But you are right, it is pretty unlikely
5.
▲
by
ahadrana
15y ago
Sorry, our github repository had some accidental check-ins that we needed to remove. I will share the link to the code shortly.
6.
▲
by
ahadrana
15y ago
Hi, I work at commoncrawl, so I will try to answer your question. We store our crawl data on S3 in the form of 100MB compressed archives and there are between 40,000 and 50,000 such files in commoncrawl’s bucket today. The key to scanning s
7.
▲
by
ahadrana
15y ago
Hi. I work for commoncrawl. We are about to start an improved recrawl and will be doing this more frequently going forward. In the process we will also consolidate our data on S3 to keep it relevant. But, as with any crawl of the Internet,