3 ms·
Caselaw data is public domain, but it's very expensive to digitize -- partly because it's mostly stored on paper, and partly because it's mixed together with co
by JackC 8y ago
Caselaw data is public domain, but it's very expensive to digitize -- partly because it's mostly stored on paper, and partly because it's mixed together with copyrighted material.
For this project we had to scan 40,000 volumes of caselaw. We used a high speed scanner at the Harvard Law Library, and went through about 40 million pages at a rate of 500,000 pages a week over a couple of years. The pages then had to be redacted of copyrighted material like headnotes inserted by private publishers, since courts typically don't publish the cases themselves, and those redactions had to be checked by humans.
That work was funded by a startup, Ravel, which is why we ended up with temporary limits on commercial use of the data. No later than March 2024, however, it will all be fully available for bulk download by anyone in the world. If necessary we'll set up a torrent. :)
(Hopefully earlier! For any state that starts officially publishing its caselaw in digital form, we can immediately release their caselaw back to the beginning, as we have already for Illinois and Arkansas.)
- extortionist 8y agoOne note that seems relevant to the grandparent comment, Ravel is now owned by LexisNexis.