6 ms·
I'm quoted in this article. Happy to discuss what we're working on at the Library Innovation Lab if anyone has questions. There's lots of people making copies
by JackC 2y ago
I'm quoted in this article. Happy to discuss what we're working on at the Library Innovation Lab if anyone has questions.
There's lots of people making copies of things right now, which is great -- Lots Of Copies Keeps Stuff Safe. It's your data, why not have a copy?
One thing I think we can contribute here as an institution is timestamping and provenance. Our copy of data.gov is made with https://github.com/harvard-lil/bag-nabit https://github.com/harvard-lil/bag-nabit , which extends BagIt format to sign archives with email/domain/document certificates. That way (once we have a public endpoint) you can make your own copy with rclone, pass it around, but still verify it hasn't been modified since we made it.
Some open questions we'd love help on --
* One is that it's hard to tell what's disappearing and what's just moving. If you do a raw comparison of snapshots, there's things like 2011-glass-buttes-exploration-and-drilling-535cf being replaced by 2011-glass-buttes-exploration-and-drilling-236cf, but it's still exactly the same data; it's a rename rather than a delete and add. We need some data munging to work out what's actually changing.
* Another is how to find the most valuable things to preserve that aren't directly linked from the catalog. If a data.gov entry links to a csv, we have it. If it links to an html landing page, we have the landing page. It would be great to do some analysis to figure out the most valuable stuff behind the landing pages.
- jszymborski 2y agoHi! Is there any one place that would be easiest for folks to grab these snapshots from? Would love to try my hand at finding documents that moved/documents that were removed.
- JackC 2y agoHmm, I can put them here for now: https://source.coop/harvard-lil/data-gov-metadata https://source.coop/harvard-lil/data-gov-metadata Unfortunately it's a bit messy because we weren't initially thinking about tracking deletions. data_20241119.jsonl.zip (301k rows) and data_20250130.jsonl.zip (305k rows) are simple captures of the API on those dates. data_db_dump_20250130.jsonl.zip (311k rows) is a sqlite dump of all the entries we saw at some point between those dates. My hunch is there's something like 4,000 false positives and 2,000 deletions between the 311k and 305k set, but that could be way off.
- jszymborski 2y agoVery cool! I take a look :)
- chrishoyle 2y agoI'd love to learn more about what is in scope of the Library Innovation Lab projects. Is it targeting data.gov specifically or all government agency websites? Given the rapid take downs of websites (cdc, usaid) do you have a prioritization framework for which website pages to prioritize or do you have "comprehensive" coverage of pages (in scope of the project)? As you allude to, I've been having a hard time learn about what sort of duplicate work might be happening given that there isn't a great "archived coverage" source of truth for government websites (between projects such as End of Term archive, Internet archive, research labs, and independent archivists). Your open questions are interesting. Content hashes for each page/resource would be a way to do quick comparisons, but I assume you might want to set some threshold to determine how much it's changed vs if it changed? Is the second question about figuring out how to prioritize valuable stuff behind two depth traversals? (ex data.gov links to another website and that website has a csv download)
- JackC 2y agoAs a library, the very high level prioritization framework is "what would patrons find useful." That's how we started with data.gov and federal Github repos as broad but principled collections; there's likely to be something in there that's useful and gets lost. Going forward I think we'll be looking for patron stories along the lines of "if you could get this couple of TB of stuff it would cover the core of what my research field depends on." In practice it's some mix of, there aren't already lots of copies, it's valuable to people, and it's achievable to preserve. > Is the second question about figuring out how to prioritize valuable stuff behind two depth traversals? Right -- how do you look at the 300,000 entries and figure out what's not at depth one, is archivable, and is worth preserving? If we started with everything it would be petabytes of raw datasets that probably shouldn't be at the top of the list.
- josh-sematic 2y agoA common metric for how much actual content has changed is the Jaccard Index. Even for large numbers of datasets that are too large to fit in memory it can be approximated with various forms of MinHash algorithms. Some write up here: https://blog.nelhage.com/post/fuzzy-dedup/ https://blog.nelhage.com/post/fuzzy-dedup/ https://en.wikipedia.org/wiki/Jaccard_index https://en.wikipedia.org/wiki/Jaccard_index
- mrshadowgoose 2y agoJust commenting to double-down on the need for cryptographic timestamping - especially in the current era of generative AI.
- _heimdall 2y agoHow does that work exactly? Does it all still hinge on trusting a know Time Stamp Authority, or is there some way of time stamping in a trustless manner?
- sebmellen 2y agoThis is the one thing blockchains are truly good for.
- _heimdall 2y agoYeah it definitely could be, though you may similarly find yourself in a spot of trusting a limited number of nodes that guarantee the chain was never tampered with.
- Retric 2y agoFor something like this there’s ways to minimize how much you need to trust nodes such as regularly publishing hashes to 3rd parties like HN. Not so useful if something was edited a few minutes after posting, but it makes it more difficult for a new administration to suddenly edit a bunch of old data.
- kergonath 2y ago> there’s ways to minimize how much you need to trust nodes such as regularly publishing hashes to 3rd parties like HN. But you could do the same thing with any hashes, right? There is no need for a blockchain in the middle.
- 2y ago
- alexvoda 2y agoTrump did this last time too. Is there a difference in the level of preparedness in archiving data compared to last time? If so, in what way is it different? Is there institutional or independent preparedness?
- JackC 2y ago(Note my lab isn't partisan and this isn't a partisan effort; public data always needs saving. But there's definitely a reason people are paying attention right now.) I think in some ways the community was less prepared this time, because there was a lot of investment in 2016-2017 and then many of the archives created at that point didn't end up being used; partly because the changes at the federal level turned out to be smaller and slower in 2017 than they're looking like this time. So some people didn't choose to invest that way this time around. [Edit: this means I think it's really important that data archives are useful. Sorting through data and putting a good interface on it should help people out today as well as being good prep for the future.] In other ways there's much more preparation; EOT Archive now has a regular practice of crawling .gov websites before and after each change of administration, which is a really great way of giving citizens a sense of how their government evolves. It will just tend to miss data that you can't click to in a generic crawl.
- _y5hn 2y agoThe P2025 timeline https://www.reddit.com/r/PrepperIntel/s/scOu1QuhNt https://www.reddit.com/r/PrepperIntel/s/scOu1QuhNt "Project Russia" is spiritual warfare https://washingtonspectator.org/project-russia-reveals-putins-playbook/ https://washingtonspectator.org/project-russia-reveals-putin... Billionaire ransacking the Treasury https://techcrunch.com/2025/02/01/senator-warns-of-national-security-risks-after-elon-musks-doge-granted-full-access-to-sensitive-treasury-systems/ https://techcrunch.com/2025/02/01/senator-warns-of-national-... Bernie's statement about this https://m.youtube.com/watch?v=mL0crkf5Dzw https://m.youtube.com/watch?v=mL0crkf5Dzw
- 1659447091 2y ago> "Project Russia" is spiritual warfare > https://washingtonspectator.org/project-russia-reveals-putin https://washingtonspectator.org/project-russia-reveals-putin... What in the Cold War conspiracy theory was that... >> The Kremlin’s design necessarily depends on the adoption of a single world belief system or religion. Expect a syncretic, gnostic blend, rooted in hierarchy — the Russian Orthodox Church at the core, and other religious factions accorded favor based on demonstrated fealty. Good luck with that. The most popular religions today are forked versions of one guys story that they couldn't agree on and have been involved in acts of genocide towards one another because of it--for centuries.
- smrtinsert 2y agoThank you for this effort.
- LastTrain 2y agoHow can people help? Sounds like a global index of sources is needed and the work to validate those sources, over time, parceled out. Without something coordinated I feel like it is futile to even jump in.
- JackC 2y agoI spent a bunch of time on this project feeling like it was futile to jump in and then just jumped in; messing with data is fun even if it turns out someone else has your data. But the government is huge; if you find an interesting report and then poke around for the .gov data catalog or directory index structure or whatever that contains it, you're likely to find a data gathering approach no one else is working on yet. There's coordinated efforts starting to come together in a bunch of places -- some on r/datahoarders, some around specific topics like climate data (EDGI) or CDC data, there's datasets being posted on archive.org. I think one way is to find a topic or kind of data that seems important and search around for who's already doing it. Eventually maybe there'll be one answer to rule them all, but maybe not; it's just so big.
- codetrotter 2y ago> sign archives with email/domain/document certificates I do a bit of web archival for fun, and have been thinking about something. Currently I save both response body and response headers and request headers for the data I save from the net. But I was thinking that maybe if instead of just saving that, I could go a level deeper and preserve actual TCP packets and TLS key exchange stuff. And then, I might be able to get a lot of data provenance “for free”. Because if in some decades when we look back at the saved TCP packets and TLS stuff, we would see that these packets were signed with a certificate chain that matches what that website was serving at the time. Assuming of course that they haven’t accidentally leaked their private keys in the meantime and that the CA hasn’t gone rogue since etc. To me I think that would make sense to build out web archival infra that preserves the CA chain and enough to be able to see later that it was valid. And if many people across the world save the right parts we don’t have to trust each other in order to verify that data that the other saved was also really sent by the website our archives say it was from. For example maybe I only archived a single page from some domain, and you saved a whole bunch of other pages from that domain around the same time so the same certificate chain was used in the responses to both of us. Then I can know that the data you are saying you archived from them really was served by their server because I have the certificate chain I saved to verify that.
- whatevermom 2y agoIt’s an interesting idea for sure. Some drawbacks I can think off: - bigger resource usage. You will need to maintain a dump of the TLS session AND an easily extractable version - difficulty of verification. OpenSSL / BoringSSL / etc. will all evolve and say, completely remove support for TLS versions, ciphers, TLS extensions… This might make many dumps unreadable in the future, or requiring the exact same version of a given software to read it. Perhaps adding the decoding binary to the dump would help, but then, you’d get Linux retro-compatibility issues. - compression issues: new compression algorithms will be discovered and could reduce data usage. You’ll have a hard time doing that since TLS streams will look random to the compression software. I don’t know. I feel like it’s a bit overkill — what are the incentives for tampering with this kind of data? Maybe a simpler way of going about it would be to build a separate system that does the « certification » after the data is dumped; combined with multiple orgs actually dumping the data (reproducibility), this should be enough the prove that a dataset is really what it claims to be.
- glitchcrab 2y agoVery tangentially related, but it always makes me smile to see rclone mentioned in the wild - its creator ncw was the CEO of the previous company I worked at.