3 ms·
Neat idea! How do you parse the PDFs?
by cerved 3y ago
Neat idea!
How do you parse the PDFs?
- petarb 3y agoApache Tika has worked well for me in the past, ended up running it on an AWS Lambda https://tika.apache.org/ https://tika.apache.org/
- nnechm 3y agoThanks, will try that.
- nnechm 3y agoPdf parsing was more tedious that I would have liked at this stage so I stuck to the SEC which requires that companies file in a text format :) so that helped. I used poppler on a digital ocean droplet, but the sheer variety of company pdfs especially european companies, some of which have to be OCRed, meant results were not really uniform. GPT still does very well, but not as well as on text documents directly. So in short, this is next on the list...
- andylynch 3y agoAre you pulling out the inline XBRL or taking a different approach? I'm curious to hear others' practical experience with it good or bad.
- nnechm 3y agoI am not using inline XBRL, instead relying on gpt's ability to read tables.
- unparagoned 3y agoHow exactly do you do that. Are you passing the raw HTML, or just the extracted text?
- yawnxyz 3y agoI have a tool that applies LLMs to abstracts and research papers — @opendocsg/pdf2md on Node / Sveltekit has been really good for me
- nnechm 3y agoThanks ! It's really interesting to see how LLMs have now brought back projects that were last updated years ago back into the limelight :)