3 ms·
There's a bunch. Here's what I do (for black-and-white text; I'm not sure how to deal with more complex scenarios): Scan with 600dpi resolution. Nevermind that
by generationP 4y ago
There's a bunch. Here's what I do (for black-and-white text; I'm not sure how to deal with more complex scenarios):
Scan with 600dpi resolution. Nevermind that this gives huge output files; you'll compress them to something much smaller at the end, and the better your resolution, the stronger compression you can use without losing readability.
While scanning, periodically clean the camera or the scanner screen, to avoid speckles of dirt on the scan.
The ideal output formats are TIF and PNG; use them if your scanner allows. PDF is also fine (you'll then have to extract the pages into TIF using pdfimages or using ScanKromsator). Use JPG only as a last resort, if nothing else works.
Once you have TIF, PNG or JPG files, put them into a folder. Make sure that the files are sorted correctly: IIRC, the numbers in their names should match their order (i.e., blob030 must be an earlier page than blah045; it doesn't matter whether the numbers are contiguous or what the non-numerical characters are). (I use the shell command mmv for convenient renaming.)
Import this folder into ScanTailor ( https://github.com/4lex4/scantailor-advanced/releases https://github.com/4lex4/scantailor-advanced/releases ), save the project, and run it through all 6 stages.
Stage 1 (Fix Orientation): Use the arrow buttons to make sure all text is upright. Use Q and W to move between pages.
Stage 2 (Split Pages): You can auto-run this using the |> button, but you should check that the result is correct. It doesn't always detect the page borders correctly. (Again, use Q and W to move between pages.)
Stage 3 (Deskew): Auto-run using |>. This is supposed to ensure that all text is correctly rotated. If some text is still skew, you can detect and fix this later.
Stage 4 (Select Content): This is about cutting out the margins. This is the most grueling and boring stage of the process. You can auto-run it using |>, but it will often cut off too much and you'll have to painstakingly fix it by hand. Alternatively (and much more quickly), set "Content Box" to "Disable" and manually cut off the most obvious parts without trying to save every single pixel. Don't worry: White space will not inflate the size of the ultimate file; it compresses well. The important thing is to cut off the black/grey parts beyond the pages. In this process, you'll often discover problems with your scan or with previous stages. You can always go back to previous stages to fix them.
Stage 5 (Margins): I auto-run this.
Stage 6 (Output): This is important to get right. The despeckling algorithm often breaks formulas (e.g., "..."s get misinterpreted as speckles and removed), so I typically uncheck "Despeckle" when scanning anything technical (it's probably fine for fiction). I also tend to uncheck "Savitzki-Golay smoothing" and "Morphological smoothing" for some reason; don't remember why (probably they broke something for me in some case). The "threshold" slider is important: Experiment with it! (Check which value makes a typical page of your book look crisp. Be mindful of pages that are paler or fatter than others. You can set it for each page separately, but most of the time it suffices to find one value for the whole book, except perhaps the cover.) Note the "Apply To..." buttons; they allow you to promote a setting from a single page to the whole book. (Keep in mind that there are two -- the second one is for the despeckling setting.)
Now look at the tab on the right of the page. You should see "Output" as the active one, but you can switch to "Fill Zones". This lets you white-out (or black-out) certain regions of the page. This is very useful if you see some speckles (or stupid write-ins, or other imperfections) that need removal. I try not to be perfectionistic: The best way to avoid large speckles is by keeping the scanner clean at the scanning stage; small ones aren't too big a deal; I often avoid this stage unless I know I got something dirty. Some kinds of speckles (particularly those that look like mathematical symbols) can be confusing in a scan.
There is also a "Picture Zones" rider for graphics and color; that's beyond my paygrade.
Auto-run stage 6 again at the end (even if you think you've done everything -- it needs to recompile the output TIFFs).
Now, go to the folder where you have saved your project, and more precisely to its "out/" subfolder. You should see a bunch of .tif files, each one corresponding to a page. Your goal is to collect them into one PDF. I usually do this as follows:
tiffcp *.tif ../combined.tif
tiff2pdf -o ../combined.pdf ../combined.tif
rm -v ../combined.tif
Thus you end up with a PDF in the folder in which your project is.
Optional: add OCR to it; add bookmarks for chapters and sections; add metadata; correct the page numbering (so that page 1 is actual page 1). I use PDF-XChangeLite for this all; but use whatever tool you know best.
At that point, your PDF isn't super-compressed (don't know how to get those), but it's reasonable (about 10MB per 200 pages), and usually the quality is almost professional.
Uploading to LibGen... well, I think they've made the UI pretty intuitive these days :)
PS. If some of this is out of date or unnecessarily complicated, I'd love to hear!
- crazygringo 4y ago> At that point, your PDF isn't super-compressed (don't know how to get those) As far as I know, it's making sure your text-only pages are monochrome (not grayscale) and to use Group4 compression for them, which is actually what fax machines use (!) and is optimized specifically for monochrome text. Both TIFF and PDF's support Group4 -- I use ImageMagick to take a scanned input page and run grayscale, contrast, Group4 monochrome encoding, and PDF conversion in one fell swoop which generates one PDF per page, and then "pdfunite" to join the pages. Works like a charm. I'm not aware of anything superior to Group4 for regular black and white text pages, but would love to know if there is.
- generationP 4y agoOh, I should have said that I scan in grayscale, but ScanTailor (at stage 6) makes the output monochrome; that's what the slider is about (it determines the boundary between what will become black and what will become white). So this isn't what I'm missing. I am not sure if the result is G4-compressed, though. Is there a quick way to tell?
- crazygringo 4y agoOn my system I can run 'pdfimages -list' on a PDF it gives me all the images in a PDF with their encoding format. The utility comes with 'poppler-utils' I believe. And I'm just now discovering by checking on my own PDF's, that 'ocrmypdf' will automatically convert Group4 to lossless JBIG2 (if optimizations are enabled) which is supposedly even more efficient for monochrome -- but encoders aren't always available [1]. I don't think ImageMagick has been updated yet to support outputting JBIG2 for PDF's. [1] https://ocrmypdf.readthedocs.io/en/latest/jbig2.html https://ocrmypdf.readthedocs.io/en/latest/jbig2.html
- homarp 4y agobeware with JBIG: "Undetectable Data Corruption in JB2/JBIG2" https://news.ycombinator.com/item?id=32537073 https://news.ycombinator.com/item?id=32537073