3 ms·
Anyone use these approaches with academic pdfs?
by sinandrei 9mo ago
Anyone use these approaches with academic pdfs?
- urschrei 9mo agoAnother approach is to teach Claude Code how to use your Zotero library's full-text search: https://github.com/urschrei/zotero_search_skill https://github.com/urschrei/zotero_search_skill.
- amelius 9mo agoAnyone using them for electronics datasheets?
- bradfa 9mo agoI would like to. I haven't yet found a solution that works well. The problems with datasheets is tables which span multiple pages, embedded images for diagrams and plots, they're generally PDFs, and only sometimes are they 2-column layout. Converting from PDF to markdown while retaining tables correctly seems to work well for me with Mistral's latest OCR model, but this isn't an open model. Using docling with different models has produced much worse results.
- sosojustdo 9mo agoI've been working on a tool specifically to handle these messy PDF-to-Markdown conversions because I ran into the same issues with tables and multi-column layouts. I’ve optimized https://markdownconverter.pro/pdf-to-markdown https://markdownconverter.pro/pdf-to-markdown to handle complex PDFs, including those tricky tables that span multiple pages and 2-column formats that usually trip up tools like Docling. It also extracts embedded diagrams/images and links them properly in the output. Full disclosure: I'm the developer behind it. I’d love to see if it handles your specific datasheets better than the models you've tried. Feel free to give it a spin!
- bradfa 9mo agoCool! But given that often electronics documentation is covered by NDAs, my preferred solution is local-first if at all possible.
- alansaber 9mo agoI've not seen any impressive products. But products do exist ie https://scibite.com/solutions/semantic-search/ https://scibite.com/solutions/semantic-search/