5 ms·
I've been fighting trying to chunk SEC filings properly, specifically surrounding the strange and inconsistent tabular formats present in company filings. This
by faxmeyourcode 2y ago
I've been fighting trying to chunk SEC filings properly, specifically surrounding the strange and inconsistent tabular formats present in company filings.
This is giving me hope that it's possible.
- otoburb 2y ago>>I've been fighting trying to chunk SEC filings properly, specifically surrounding the strange and inconsistent tabular formats present in company filings. For this specific use case you can also try edgartools[1] which is a library that was relatively recently released that ingests SEC submissions and filings. They don't use OCR but (from what I can tell) directly parse the XBRL documents submitted by companies and stored in EDGAR, if they exist. [1] https://github.com/dgunning/edgartools https://github.com/dgunning/edgartools
- faxmeyourcode 2y agoI'll definitely be looking into this, thanks for the recommendation! Been playing around with it this afternoon and it's very promising.
- barrenko 2y agoIf you'd kindly tl;dr the chunking strategies you have tried and what works best, I'd love to hear.
- anirudhb99 2y ago(from the gemini team) we're working on it! semantic chunking & extraction will definitely be possible in the coming months.
- jgalt212 2y agoisn't everyone on iXBRL now? Or are you struggling with historical filings?
- faxmeyourcode 2y agoXBRL is what I'm using currently, but it's still kind of a mess (maybe I'm just bad at it) for some of the non-standard information that isn't properly tagged.