2 ms·
So we're storing non-compressed 300dpi color scans of every page, as well as the original shrinkwrapped books out in a saltmine somewhere -- we're not losing an
by JackC 11y ago
So we're storing non-compressed 300dpi color scans of every page, as well as the original shrinkwrapped books out in a saltmine somewhere -- we're not losing any data.
There's a whole separate problem of turning those scans into a high-quality data set. The first pass will be decent-but-not-perfect OCR of the full text (with page-location data, like Google Books), plus human-checked metadata for stuff like case name, judge, and date. Since we have the original scans as well, there's lots of room to iteratively improve the data conversion from there via ReCAPTCHA and the like.