3 ms·
Really great work! Two questions: What sort OCR stack did you use? Is there a way to see the text inside the search results? I'm only seeing the PDFs themsel
by bpchaps 9y ago
Really great work!
Two questions:
What sort OCR stack did you use?
Is there a way to see the text inside the search results? I'm only seeing the PDFs themselves and would love to do some full text searches of my own!
- jiscariot 9y agoThanks much for the feedback! Imagemagick -> tesseract -> solr/lucene I am a neophyte when it comes to this stuff, so I'm sure someone with more experience could get better results from tweaking IM/tess. Some of the IM convert stuff was extremely memory intensive on larger documents and AWS was starting to get really expensive. Later on I added PDFbox to split the PDFs pre IM and run a page at a time vs. the entire document. SOLR has a highlighting feature that I never really got working right. That would have showed some context to the search terms in the results.