2 ms·
Letting Claude work a little longer produced this behemoth of a script (which is supposed to be somewhat universal in correcting similar OCR'd PDFs - not yet te
by dperfect 8mo ago
Letting Claude work a little longer produced this behemoth of a script (which is supposed to be somewhat universal in correcting similar OCR'd PDFs - not yet tested on any others though):
https://pastebin.com/PsaFhSP1 https://pastebin.com/PsaFhSP1
which uses this Rust zlib stream fixer:
https://pastebin.com/iy69HWXC https://pastebin.com/iy69HWXC
and gives the best output I've seen it produce:
https://imgur.com/itYWblh https://imgur.com/itYWblh
This is using the same OCR'd text posted by commenter Joe.
- daveguy 8mo ago> which is supposed to be somewhat universal in correcting similar OCR'd PDFs Xerox would like a word. https://news.ycombinator.com/item?id=29223815 https://news.ycombinator.com/item?id=29223815 Point being, "correcting" to "correct looking" may be worse than just accepting errors. Errors are often clearly identified by humans as a nonsense word. "Correcting" OCR can result in plausible, but wrong results that are more difficult for the human in the loop to identify.
- dperfect 8mo agoThat's true if we're correcting OCR of actual output text. In this case, it's operating on the base 64 text, trying to produce chunks that form valid zlib streams and PDF syntax so the file can be intact enough to be opened. "Just accepting errors" would mean not seeing any content in the file because it cannot be read. So yes, the "fixed" output has errors, but it’s not hallucinating details like an LLM, nor is it trying to produce output that conforms to any linguistic or stylistic heuristics. The phrase "correcting similar OCR'd PDFs" should have been "correcting similar OCR'd base 64 representations of PDFs".