Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
bnewbold
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
bnewbold
5y ago
The author of this post (Mako) is standing next to me right now, and was initially excited that he could comment here on HN without creating an account. But of course the "add comment" form directs to a signup page!
32.
▲
by
bnewbold
6y ago
Great question! BASE, SHARE ( https://share.osf.io/ ), and CORE ( https://core.ac.uk ) all primarily pull metadata via OAI-PMH, though they may also incorporate other sources these days. We have worked with CORE to
33.
▲
by
bnewbold
6y ago
This is great feedback, thank you. For future follow-up, my work email is my handle here (bnewbold) at archive.org
34.
▲
by
bnewbold
6y ago
I think it is in a good place for simple bibliometric queries. The fatcat elasticsearch API is open at https://api.fatcat.wiki/fatcat_release/ (behind a proxy to filter "unsafe" requests). That works pretty w
35.
▲
by
bnewbold
6y ago
Ah, sorry to hear. We in particular want to include content from outside the US/Europe publishing world. For Japanese publishing, we have done metadata imports from JaLC (Japanese DOI registrar), and crawled a lot of open content from
36.
▲
by
bnewbold
6y ago
Fixed, thanks!
37.
▲
by
bnewbold
6y ago
We are mostly not indexing on a journal-by-journal basis, but try to import from large, broad sources. For example, DOI registrars (Crossref, Datacite, J-Stage), DOAJ article and journal metadata (for OA publications), etc. Some field-speci
38.
▲
by
bnewbold
6y ago
Thank you for the kind words! We are friendly with Semantic Scholar, and have used their "open corpus" dumps as one of several URL seed lists for crawling in the past. Their search and discovery tech is more sophisticated than our
39.
▲
by
bnewbold
6y ago
This service was hinted at back in September, but is now formally announced and live at https://scholar.archive.org Related previous post: https://news.ycombinator.com/item?id=24485444 Much of the catalog functi
40.
▲
Internet Archive Scholar: Search Millions of Research Papers
(blog.archive.org)
342 points
by
bnewbold
6y ago
|
47 comments
41.
▲
by
bnewbold
6y ago
At the Internet Archive, we are working on one aspect of this problem: https://fatcat.wiki/ Other notable efforts, mostly envisioned and led by librarians, are "dark" digital archives (LOCKSS, CLOCKSS, Portico, et
42.
▲
by
bnewbold
6y ago
There are APIs and it would be great if more people and organizations built on top of them, and specifically build content or collection-specific interfaces. Here is the entry point for API documentation: https://archive.org&#x
43.
▲
by
bnewbold
6y ago
Performance is fun! One aspect is that our data centers are in California, with no CDN. If you are on the other side of the world, you will have higher round-trip latency on every request, for all services. Another is layers of caching. Pop
44.
▲
by
bnewbold
6y ago
The disks are spinning all the time, and most disks are seeing fairly frequent reads to some content or another. A lot of content is very rarely accesses, but almost every disk has some content which gets accessed. If spinning disks had onl
45.
▲
by
bnewbold
6y ago
If you are interested in cost modeling for long-term digital preservation, check out this blog series: https://blog.dshr.org/2019/02/economic-models-of-long-term-s...
46.
▲
by
bnewbold
6y ago
Link back to recent discussion on this topic: https://news.ycombinator.com/item?id=24422593
47.
▲
by
bnewbold
6y ago
From what I have seen, the least technically resourced journals often use hosted platforms or free software like OJS (basically wordpress for journals), which comes with features like HTML meta tags and OAI-PMH by default. The trickier case
48.
▲
by
bnewbold
6y ago
Unpaywall is very helpful! However, even for direct PDF links, publishing platforms will often do things like check for a session cookie; if you don't have the correct cookie you get bounced back to the landing page, where you need fin
49.
▲
by
bnewbold
6y ago
In Latin America, the SciELO network has been very successful at providing shared, low-cost, stable infrastructure for digital journal hosting using state funding: https://en.wikipedia.org/wiki/SciELO
50.
▲
by
bnewbold
6y ago
At the Internet Archive, we are working on this exact problem, and have been in communication with the pre-print's authors. We have built open infrastructure (open source, open data) tracking "preservation coverage", for exam
51.
▲
by
bnewbold
8y ago
Have you looked at the WARC format? It's ridiculously simple, basically concatenated raw HTTP requests and responses, with some extra HTTP metadata headers mixed in (a la extra JSON metadata keys). You can open it with a text editor. V
52.
▲
by
bnewbold
8y ago
The Archive currently has about 46 Petabytes of content ("bytes archived"), and over 120 PB of raw disk capacity; the difference is due to data replication, "currently filling" storage, non-storage infrastructure, etc. W
53.
▲
by
bnewbold
8y ago
They are very different beasts. Arxiv (and most pre-print repos) accept submissions (with filters, like requiring academic affiliation or vouching, and requiring reasonably-formatted metadata), have a moderator do a skim-level review of the
54.
▲
by
bnewbold
8y ago
I could be misinterpreting your comment, but it reads like a broad misunderstanding of the role, economics, and value-add of contemporary publishing. Here is a (now somewhat old) breakdown of per-article costs, contrasting open access and s
55.
▲
by
bnewbold
8y ago
I can't speak formally for the Internet Archive, but the existing content and services are not going to disappear overnight: funding comes from several sources, thought has been put in to organizational structure, and things have been
56.
▲
Catching the Wave: The Tide Turns Toward the Subscription Model
(scholarlykitchen.sspnet.org)
1 points
by
bnewbold
9y ago
|
0 comments
57.
▲
by
bnewbold
9y ago
Well, I oversimplified a bit. The "tarball" (WARC or ARC file) is a single large file with individually compressed HTTP response objects concatenated together. The index stores the byte offset and length (compressed) into this fil
58.
▲
by
bnewbold
9y ago
The Wayback backend is much simpler than one might imagine. The index lookup ("where is the data for this URL at this timestamp?") is called the CDX API and is pretty fast, given that it's basically looking up a line in a sor
59.
▲
by
bnewbold
9y ago
To state the obvious, this proposal is antithetical to the concept of network neutrality. Also, the belief that there would be "competitors in the user's area that don't play stupid games" that offer comparable services
60.
▲
by
bnewbold
9y ago
Internet Archive | San Francisco, CA | PM, SRE, Book Curator | Full Time, ONSITE The Internet Archive is a US-based non-profit which has been backing up the web and pursuing "Universal Access to All Knowledge" for more than 20 yea
More ›