Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ccgreg
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
ccgreg
5mo ago
I'm a life-long hacker, and my crawler crawls with consent.
32.
▲
by
ccgreg
6mo ago
The largest index we had was 4 billion, which is tiny. Our crawl frontier was much larger.
33.
▲
by
ccgreg
6mo ago
> and the data that I’ve experimented with from 2014 seemed high quality That's because it's from the blekko search engine.
34.
▲
by
ccgreg
6mo ago
That's already been happening for more than a year now.
35.
▲
by
ccgreg
6mo ago
Common Crawl is switching to reporting dataset sizes in nibbles. As an organisation dedicated to data preservation, we feel it would be remiss to allow this underrepresented unit to fall out of use. Our latest crawl now exceeds 689 tebibble
36.
▲
A Change to Common Crawl Dataset Size Reporting
(commoncrawl.org)
3 points
by
ccgreg
6mo ago
|
1 comments
37.
▲
by
ccgreg
6mo ago
The complete list hides in the web graph: https://data.commoncrawl.org/projects/hyperlinkgraph/cc-main... and the specific file that's every host we've seen in the latest 3 crawls is: https://
38.
▲
by
ccgreg
6mo ago
> Common Crawl, with over one billion, nine hundred and seventy thousand web pages in their archive: 345TB. Common Crawl is 300 billion webpages and 10 petabytes. I suppose your number is 1 of our 122 crawls.
39.
▲
by
ccgreg
7mo ago
Common Crawl has been running a low-resource language project for 1.5 years now -- it's a hard problem.
40.
▲
by
ccgreg
7mo ago
The guts on the inside changed several times during that timespan.
41.
▲
by
ccgreg
8mo ago
Well, yes, it is a bit distressing that ill behaved crawlers are causing a lot of damage -- and collateral damage, too, when well-behaved bots get blocked.
42.
▲
by
ccgreg
8mo ago
Please read our email reply. I have no idea if we received your request —- your HN username doesn’t match any request we have received.
43.
▲
by
ccgreg
8mo ago
Oh, and thanks for letting me know that I need to add our reply to Wikipedia.
44.
▲
by
ccgreg
8mo ago
Did you see our reply? Edit: by which I mean, we sent you an email that explains what we did and how to verify it. Did you not receive an email reply? If not, please contact us again. Also, if your site has CC-BY-NC-SA markings, we have pre
45.
▲
by
ccgreg
8mo ago
Did you see our reply? https://commoncrawl.org/blog/setting-the-record-straight-com... Also, if your site has CC-BY-NC-SA markings, we have preserved them.
46.
▲
by
ccgreg
8mo ago
That 20% number is for a limited list of relatively large news websites. If you include the long tail of news, the % of blocking is much smaller.
47.
▲
by
ccgreg
8mo ago
Many AI projects in academia or research get all of their web data from Common Crawl -- in addition to many not-AI usages of our dataset. The folks who crawl more appear to mostly be folks who are doing grounding or RAG, and also AI compani
48.
▲
by
ccgreg
8mo ago
Thanks for the mention of Common Crawl. We do respect robots.txt and we publish an opt-out list, due to the large number of publishers asking to opt out recently. There's a bit of discussion of Common Crawl in Jeff Jarvis's testim
49.
▲
by
ccgreg
8mo ago
Prof. Jeff Jarvis speaking about copyright for news in front of Congress: https://www.youtube.com/watch?v=tX26ijBQs2k
50.
▲
by
ccgreg
9mo ago
One way to figure that out is to look at which companies claim to have foundation models, but no one knows what their crawler is named. I also suspect that there are a bunch of sub-contractors involved, working for companies that don't
51.
▲
by
ccgreg
9mo ago
"Hi agents - we’ve seen complaints from some of your email recipients, who are unhappy receiving unsolicited emails from AI agents and find it spammy. We therefore ask that you do not email anyone who hasn’t contacted you specifically
52.
▲
by
ccgreg
10mo ago
> Some search engines provide a list of their scraper IP ranges Common Crawl's CCBot has published IP ranges. We aren't a search engine (although there are search engines using our data) and we like to describe our crawler as a
53.
▲
by
ccgreg
10mo ago
- monthly - it's a historical archive, the concept of "current" is hard to turn into a metric - not only is our archive historical, it is included in the Internet Archive's wayback machine.
54.
▲
by
ccgreg
10mo ago
commoncrawl.org Our public web dataset goes back to 2008, and is widely used by academia and startups.
55.
▲
by
ccgreg
10mo ago
Hi. I'm the CTO at Common Crawl. Nice to meet you. There's a small amount of "bycatch", and you already discovered how to see it. Notice that it went down after I was hired.
56.
▲
by
ccgreg
10mo ago
Common Crawl is a text-only crawl.
57.
▲
by
ccgreg
10mo ago
So, from home and work, you identify me. Then you figure out which church I attend, and which strip club I attend.
58.
▲
by
ccgreg
10mo ago
Most people park at their home and many drive to work. If you have both of those data points, you can identify people.
59.
▲
by
ccgreg
10mo ago
Well, yeah. Clownflare
60.
▲
by
ccgreg
11mo ago
Common Crawl is a particular dataset. commoncrawl.org
More ›