4 ms·
I work at the Harvard Library Innovation Lab with the folks who are making this happen. Super excited that it's finally public. If anyone has questions I'm happ
by JackC 11y ago
I work at the Harvard Library Innovation Lab with the folks who are making this happen. Super excited that it's finally public. If anyone has questions I'm happy to dig up answers.
Here's how big this is: we don't even know yet how many cases we'll end up with, to within the nearest million.
PSA: we're hiring a devops engineer[1]. In addition to building amazing tools to access all this data, we're running a distributed linkrot preservation service[2] that after just two years is in use by 40% of American law schools and 10% of state supreme courts; an open-textbook-as-forkable-playlist[3] tool in use at Harvard Law and a half dozen other law schools; and a research project on distributed encrypted library archives[4] for preserving high-value cultural records. We're basically the alien in the brain of a 200-year-old library -- it's a fun place to work.
[1] http://librarylab.law.harvard.edu/blog/2015/10/20/hiring-devops-energy-wanted/ http://librarylab.law.harvard.edu/blog/2015/10/20/hiring-dev...
[2] http://perma.cc http://perma.cc
[3] http://librarylab.law.harvard.edu/projects/h2o http://librarylab.law.harvard.edu/projects/h2o
[4] http://librarylab.law.harvard.edu/projects/time-capsule-encryption http://librarylab.law.harvard.edu/projects/time-capsule-encr...
- nekopa 11y agoThanks for stopping by! Do you happen to know what type of license the materials will be under when they open up access? I am an English teacher, and about 80% of my students are professional lawyers, so I wonder how free I will be to use this material in my classes. I already use Harvard Law School's free case studies, they're great, and under CC license if I remember correctly.
- JackC 11y agoBottom line: everything becomes public domain after eight years. Before then, you'll be able to search/view/download up to 500 cases per day through either a web interface or API. As far as I know there's no licensing on the individual cases you download. For academic researchers, before the eight years are up, we can also provide a full data dump -- you just have to sign an agreement not to redistribute bulk data.
- deleted 11y ago[deleted]
- wahsd 11y agoIs anyone considering if there are metadata issues that may be lost somehow through this process and that should be included through some sort of coding process? I don't have an answer, but I am wondering essentially if there is a deliberate effort to circle the question "Are we missing anything?"
- JackC 11y agoSo we're storing non-compressed 300dpi color scans of every page, as well as the original shrinkwrapped books out in a saltmine somewhere -- we're not losing any data. There's a whole separate problem of turning those scans into a high-quality data set. The first pass will be decent-but-not-perfect OCR of the full text (with page-location data, like Google Books), plus human-checked metadata for stuff like case name, judge, and date. Since we have the original scans as well, there's lots of room to iteratively improve the data conversion from there via ReCAPTCHA and the like.