Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ccgreg
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
61.
▲
by
ccgreg
11mo ago
> it used a stack machine Do you have some reading for this? I've used that compiler but I never read the resulting assembly language.
62.
▲
by
ccgreg
11mo ago
Most academic AI research and AI startups find Common Crawl adequate for what they're doing. Common Crawl also has a lot of not-AI usage.
63.
▲
by
ccgreg
1y ago
I had fun reading this -- the radio astronomy technique called VLBI (very long baseline interferometry) has a ton of overlap, but the jargon words are fairly different. There are plenty of differences: our telescopes are mostly on the surfa
64.
▲
by
ccgreg
1y ago
I've seen a Nobel Laureate disinvited from speaking at a conference due to an incredibly offensive statement. Toxic speech has consequences.
65.
▲
by
ccgreg
1y ago
> they were actually very performant Insanely expensive for that performance. I was the architect of HPC clusters in that era, and Itanic never made it to the top for price per performance. Also, having lived through the software stack i
66.
▲
by
ccgreg
1y ago
A lot of publishers do not care about blind people, and would prefer that they be unable to use AI to read.
67.
▲
by
ccgreg
1y ago
The IETF AI-Preference standard group is currently discussing whether or not to include an example of bypassing AI preferences to support assistive technologies. Oddly enough, many publishers oppose that.
68.
▲
by
ccgreg
1y ago
My experience is that a news crawl is not a big expense at scale, but so far I've only built one and inherited one. BTW No one uses blog pings, the latest hotness is IndexNow.
69.
▲
by
ccgreg
1y ago
That's not how search engines work. They have a good idea of which pages might be frequently updated. That's how "news search" works, and even small startup search engines like blekko had news search.
70.
▲
by
ccgreg
1y ago
I don't think you're correct about Google. Caching webpages is bread-and-butter for search engines, that's how they show snippets.
71.
▲
by
ccgreg
1y ago
The Fastly report[1] has a couple of great quotes that mention Common Crawl's CCBot: > Our observations also highlight the vital role of open data initiatives like Common Crawl. Unlike commercial crawlers, Common Crawl makes its dat
72.
▲
by
ccgreg
1y ago
The blekko search engine index was only 1 billion pages, compared to Common Crawl Foundation's crawl of 3 billion webpages per month.
73.
▲
by
ccgreg
1y ago
Salt water without oxygen and salt water with oxygen are different.
74.
▲
by
ccgreg
1y ago
It's a similar loophole as public libraries. When I was a kid, I read thousands of books from the library, without paying anyone anything. But as for the crawl loophole: CCBot obeys robots.txt, and CCBot also preserves all robots.txt a
75.
▲
by
ccgreg
1y ago
One way that Cloudflare is gatekeeping is by declaring which bots are AI Bots. Common Crawl's CCBot is used for a lot of stuff -- it's an archive, there are more than 10,000 research papers citing common crawl, mostly not AI -- bu
76.
▲
by
ccgreg
1y ago
As the Cloudflare post indicates, most crawlers can be verified by IP address.
77.
▲
by
ccgreg
1y ago
Conventional crawlers already have a way to identify themselves, via a json file containing a list of IP addresses. Cloudflare is fully aware of this defacto standard.
78.
▲
by
ccgreg
1y ago
Given that this discussion was started by someone mentioning only these 3 related things, I'd guess that the motivation might a negative one.
79.
▲
by
ccgreg
1y ago
You can leave off the laser beams and just look at them with telescopes. The paper we're discussing is exactly that.
80.
▲
by
ccgreg
1y ago
Undetected means you can compute an upper bound, based on the performance of your instrument. Again, these are undergraduate level concepts.
81.
▲
by
ccgreg
1y ago
No. Undetected iron and zero iron are different. Source: astronomer.
82.
▲
by
ccgreg
1y ago
> We report detection of CN emission and also detect numerous Ni I lines while Fe I remains undetected, potentially implying efficiently released gas-phase Ni. Where does it say there is zero iron? This is an upper bound, not zero.
83.
▲
by
ccgreg
1y ago
https://en.wikipedia.org/wiki/Global_dimming Sadly misunderstood by a bunch of people.
84.
▲
by
ccgreg
1y ago
Hopefully you read all of the links in the article -- the purpose of thecoversation is to present information to the general public, with references to research that the author has been involved with.
85.
▲
by
ccgreg
1y ago
> Publishing a crawl, or the URL's, under CC-0, CC-by, BSD, or Apache would make them usable without restrictions or any further legal analyses. This isn't true, and I can't imagine that any lawyer would agree with this st
86.
▲
by
ccgreg
1y ago
Historically, the Sparc 6400 was derided for not being NUMA, but instead being Uniformly Slow.
87.
▲
by
ccgreg
1y ago
> The long and short of it is that if you’re building a HPC application, or are sensitive to throughput and latency on your cutting-edge/high-traffic system design, then you need to manually pin your workloads for optimal performanc
88.
▲
by
ccgreg
1y ago
The goo.gl URLs that are publicly known are already in the Internet Archive and Common Crawl crawls.
89.
▲
by
ccgreg
1y ago
Common Crawl doesn't own the content in its crawl, so no, our terms of use do not grant anyone permission to ignore the actual content owner's license. We carefully preserve robots.txt permissions in robots.txt, in http headers, a
90.
▲
by
ccgreg
1y ago
I'm all for that.
More ›