Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ccgreg
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
91.
▲
by
ccgreg
1y ago
The team that runs the Common Crawl Foundation is well aware of how to crawl and index the web in real time. It's expensive, and it's not our mission. There are multiple companies that are using our crawl data and our web graph me
92.
▲
by
ccgreg
1y ago
At the end, the author thinks about adding Common Crawl data. Our ranking information, generated from our web graph, would probably be a big help in picking which pages to crawl. I love seeing the worked out example at scale -- I'm sur
93.
▲
by
ccgreg
1y ago
Note that a book author cannot publish a book and then refuse to let libraries buy copies and lend them out. This was litigated 100+ years ago.
94.
▲
by
ccgreg
1y ago
I don't think Reddit pays the people who voluntarily write Reddit content. Valuable to Reddit, I guess.
95.
▲
by
ccgreg
1y ago
See https://digitalcorpora.org/corpora/file-corpora/cc-main-2021... for a set of 8 million PDF files from the web, as seen by a single crawl of Common Crawl.
96.
▲
by
ccgreg
1y ago
Common Crawl's count of unique goo.gl links is approximately 10 million. That's in our permanent archive, so you'll be able to consult them in the future. No search engine or crawler person will ever recommend using a shorten
97.
▲
by
ccgreg
1y ago
GFS outputs are also available for free from NOMADS. That was true long before AWS came up with the public dataset program.
98.
▲
by
ccgreg
1y ago
We have numbers with a wide error bar.
99.
▲
Show HN: Help improve language coverage in Common Crawl
8 points
by
ccgreg
1y ago
|
0 comments
100.
▲
by
ccgreg
1y ago
In 2020 I toured the machine room and those boxes were powered off.
101.
▲
by
ccgreg
1y ago
I aggregate weather forecasts for the Event Horizon Telescope Collaboration, which is the collaboration behind those black hole images you might have seen. We want to pick the best nights during an observing window based on the weather in 1
102.
▲
by
ccgreg
1y ago
In my most recent trip through academic astronomy, not only do they say "visible light" early and often, radio astronomers refer to "optics" and "photons" and VLBI images are called "images" and not &
103.
▲
by
ccgreg
1y ago
Take a photo of the sun setting behind clouds, and then marvel that the camera still sees a big red Sun, when your eye barely sees it. That's because the camera goes way farther into the red than your eye, and the clouds let that sub-r
104.
▲
by
ccgreg
2y ago
Does that indicate the robots.txt is how "no crawl" is indicated? robots.txt doesn't have "no crawl", it has allow and disallow.
105.
▲
by
ccgreg
2y ago
robots.txt has a maximum relevant size of 500 kib.
106.
▲
by
ccgreg
2y ago
I'd say a bigger problem is that people disagree about the meaning of nofollow and noindex.
107.
▲
by
ccgreg
2y ago
How does that second paragraph work? I run engineering at Common Crawl, and Common Crawl is ethical and has never fired a developer over ethics. During the End of Term 2024 crawl[1], we discovered a lot of blocking on US government websites
108.
▲
by
ccgreg
2y ago
The best part about "verified crawlers" is that there's no easy way to discover how to become one. Or if you need to become one.
109.
▲
by
ccgreg
2y ago
Cloudflare, apparently.
110.
▲
by
ccgreg
2y ago
Many non-AI and AI research projects and companies do use Common Crawl. There's apparently only a small list of AI companies who don't.
111.
▲
by
ccgreg
2y ago
Cloudflare's documentation says that Labyrinth is not based on robots.txt.
112.
▲
by
ccgreg
2y ago
Cloudflare blocks Common Crawl's bot no matter how slow it goes.
113.
▲
by
ccgreg
2y ago
"This will waste the time of people using screen-readers, but that is a sacrifice I am willing to make."
114.
▲
by
ccgreg
2y ago
1. Find many examples of these nofollow links 2. Create a webpage with these links, not including the nofollow 3. ... 4. Profit!
115.
▲
by
ccgreg
2y ago
Common Crawl's latest crawl was Jan 12th-25th, and the index is available.
116.
▲
by
ccgreg
2y ago
Common Crawl Foundation | REMOTE | Full and part-time | https://commoncrawl.org/ | web datasets I'm the CTO at the Common Crawl Foundation, which has a 17 year old, 9 petabyte crawl & archive of the web. Our open d
117.
▲
by
ccgreg
2y ago
Crawl budget is relevant to every site in Common Crawl.
118.
▲
Publishers Target Common Crawl in Fight over AI Training Data
(wired.com)
4 points
by
ccgreg
2y ago
|
1 comments
119.
▲
by
ccgreg
2y ago
> Although Common Crawl has been essential to the development of many text-based generative AI tools, it was not designed with AI in mind. Founded in 2007, the San Francisco–based organization was best known prior to the AI boom for its
120.
▲
by
ccgreg
2y ago
Thank you!
More ›