4 ms·
There are some applications that will do a full text search of pdfs across directories, but seem geared towards server rather than desktop, with commensurate le
by crieff 10y ago
There are some applications that will do a full text search of pdfs across directories, but seem geared towards server rather than desktop, with commensurate levels of cost and complexity. Conceptually you could use image magik and lucene to make a Linux solution, but without any added features such as summary or title.
I am experimenting with a lightweight solution, but am working out which compromises are reasonable to take so that it is worthwhile but not overwhelming of the machine it runs on. Still have to give it a real test with a large number of files as well.
- cr0sh 10y agoAfter I posted, I did some searching, and it appears like something could be made using SOLR or Elasticsearch. Both seem to have methods/plugins for filesystem indexing and document importing/analysis, as well as easy interfaces to allow for any language to be used for development. Combining all of that, plus some dev work and such a search appliance looks doable for a home system, using only a single node. For the hardware, I figure I could potentially use some old stuff I have (thinking like a Core2 Quad with 16gb RAM and a large hard drive would be fine). I could probably stuff it into an old half-depth 1u server case. The problem now is finding the time to build it...
- crieff 10y agoThanks, I had missed SOLR and TIKA even though I had investigated Lucene. One criterion I had for a lightweight solution was to not require Java. No problem with Java, just that it is a big dependency and my perception is that it is not a common install on the laptop or desktop of people reading pdfs, at least out side of the STEM stream.