5 ms·
Any recommendations for converting books into another format besides PDF? Ever since I read SICP as a texinfo in Emacs while working on another screen I've been
by chachalarue 11y ago
Any recommendations for converting books into another format besides PDF? Ever since I read SICP as a texinfo in Emacs while working on another screen I've been looking for an easier automated way to convert my library to texinfo or LaTeX source.
- vitovito 11y agoIf by "automated" you mean "cheap/free," no. If by "automated" you mean "I don't have to do any work" irrespective of cost, yes. When you "scan" a book, you're taking photos of the pages, which you can then run through OCR, but OCR, even of a page scanned with a flatbed scanner, is not going to understand page layout and a variety of typefaces perfectly. You're going to have a lot of errors to correct, and usually some reformatting to do, and adding in the text that OCR missed because it was part of an image or something. You can do it yourself, comparing the scan to the OCR'd text, a layperson can edit and correct ~18 lines per minute, probably an entire weekend of your time. You can hire a professional editor, who can edit and correct ~25 lines per minute, probably a few to several hundred dollars (USD) per book. You could Mechanical Turk it, but I'm not sure the math and time trade-off works out, given that you have to have the edits confirmed and redone, either by other turkers, or by you. At the end, you could have an hOCR file that you could readily turn into an ePub or something else, but there's no magic solution. (These figures are based on research and testing I did in 2013, using hOCR editing tools and hiring and timing a professional editor versus myself.)
- wodenokoto 11y agoIf the book is just text and headings, can't you expect OCR to do a good job without any real need of human intervention?
- vitovito 11y agoDepends on what you mean by a "good job" and what you're doing with the results. 80-90% recognition isn't good enough if it means when you search for a term it doesn't show up because the OCR saw "rn" and wrote "m", or if you're having it translated, or having it read aloud with a text-to-speech synthesizer. In my tests, we were seeing accuracy problems of 10-15% of lines needing correction, and this is a book that was primarily headers and text. Sometimes this is character-level issues, like I cited above. Sometimes this is dust, debris, shadows or markings being confused with text. You get a little closer by running spelling and context checks against the words, but it's never 100% accurate. And if you aren't looking at the original pages, or you need automated systems to search/parse/translate/etc. the text, you need it to be 100% accurate, which means you need a human editor. OCR isn't a solved problem.
- wodenokoto 11y agoThank you for the detailed response. I honestly thought it was a solved problem (as in better performing than a human) as long as there is just running text.