5 ms·
Definitely! I'll sleep better when we have all the bulk data shipped to everyone who wants it. That will happen by March 2024, or earlier for any states that sw
by JackC 8y ago
Definitely! I'll sleep better when we have all the bulk data shipped to everyone who wants it. That will happen by March 2024, or earlier for any states that switch to official digital publishing.
(I mean, Harvard isn't a bad home for this -- I work in a building with books that predate the printing press, and I work on stuff like Creative Commons-licensed forkable textbooks. Libraries are cool places. But Harvard definitely shouldn't be the only place that preserves this data set.)
As far as preservation-friendly formats, our bulk data download format is xzipped jsonlines, which is tuned for NLP (highly compressed, parseable in a few lines of python with low memory requirements) rather than preservation:
https://case.law/bulk/download/ https://case.law/bulk/download/
Internally we have a preservation format where each volume is stored as a bagit bag containing METS XML for the OCR and case-level data, plus color and black and white images of each case. These are much harder to work with, so it's not a focus to share them right now, but we can definitely share if someone makes a case for it.