3 ms·
Oh wow! I've worked on turning PAIP (Paradigms of Artificial Intelligence Programming) from a book into a bunch of Markdown files, but that's "only" about a tho
by pronoiac 2y ago
Oh wow! I've worked on turning PAIP (Paradigms of Artificial Intelligence Programming) from a book into a bunch of Markdown files, but that's "only" about a thousand pages long, compared to the roughly 27000 pages long of all those volumes. I have advice, possibly helpful, possibly not.
Getting higher quality scans could save you some headaches. Check the Internet Archive. Or, get library copies, and the right camera setup.
Scantailor might help; it lets you semi-automate a chunk of things, with interactive adjustments. I don't know how its deskewing would compare to ImageMagick. The signature marks might be filtered out here.
I wrote out some of my process for handling scans here - https://github.com/norvig/paip-lisp/releases/tag/v1.2 https://github.com/norvig/paip-lisp/releases/tag/v1.2 . I maybe should blog about it.
If you get to the point of collaborative proofreading, I highly recommend Semantic Linefeeds - each sentence gets its own line. https://rhodesmill.org/brandon/2012/one-sentence-per-line/ https://rhodesmill.org/brandon/2012/one-sentence-per-line/ I got there by:
* giving each paragraph its own line
* then, linefeed at punctuation, maybe with quotation marks and parentheses? It's been a while
- bambax 2y agoYou are right that the quality of the scans is paramount! Unfortunately I don't have access to the physical books and have to work with the scans as they are (they're not good). But I will look at Scantailor, it looks interesting. For now I reconstruct paragraphs in html but I could do markdown just as well (where paragraph breaks are marked by double line breaks, and single line breaks don't count). Collaborative proofreading would be cool but it would require some way of properly tracking who wrote what, and I'm not sure what to use or if I should build a simple system from scratch. Do you have recommendations?
- jfil 2y agoBecause you're creating webpages from the text, one option for collaborative notes/corrections is to use a Web Annotation system like Hypothes.is.
- pronoiac 2y agoI got a copy of the 30-year old book from EBay or Amazon for $20, chopped the spine off, and fed it through a scanner. Doing that to a century-old book feels wrong! ScanTailor was tricky to start with; dunno if there's a manual. I remember belatedly realizing that there's automation at each step, that one can then quickly skim and manually adjust. For collaborative editing, git via GitHub worked for us. Tracking who did what, and when, is easy. It allowed for sweeping edits covering multiple chapters. Building some porcelain on top of that, for less technical folks, could be good.
- pronoiac 2y ago> Pour obtenir un document de Gallica en haute définition, contacter utilisation.commerciale@bnf.fr. roughly: > To obtain a Gallica document in high definition, contact utilisation.commerciale@bnf.fr. My expectations would be very low, but I'd reach out to them anyway.
- 2Gkashmiri 2y agoA few years ago I got so good at the whole scan>scantailor>PDF that I could scan a 100-150 page book, send that to scantailor, edit it and improve it to TIFF. Convert to PDF and OCR it in half hour. I got very good at this but page turning way a bore. The PDF turned out in a mechanical fashion without much effort. I made a few scripts to do TIFF to PDF and then stictching them and doing OCR.
- pronoiac 2y agoPage turning? So, non-destructive, with cameras? How's the quality?
- 2Gkashmiri 2y agoI did two things. 1. Destructive removing page staples and then just page turn manually. Or 2. Book holder for non destructive page turning. More delicate. Slow turning. CAmara I use a cheap https://www.amazon.in/ClickScan-Foldable-Metallic-Design-2-Port/dp/B0DF318WJC https://www.amazon.in/ClickScan-Foldable-Metallic-Design-2-P...