6 ms·
When you say "regardless of structure" -- if it's a relational database, that inherently implies a set structure. Or did you mean the information is consistent
by IceMetalPunk 4y ago
When you say "regardless of structure" -- if it's a relational database, that inherently implies a set structure. Or did you mean the information is consistent enough to be represented in one relational structure, but is presented in the PDFs in different formats?
- chaps 4y agoHonestly, I don't know what it would look like, but being able to query would be deeply important. But what I can say is that a lot of the PDFs I work with are auto-generated as PDF forms using queries, after the information was likely inserted into a relational database during some esoteric transcription step. Meaning, the information that formed the PDFs very likely come from a relational database and an inversion back to its original relational form is probably the convenient form. Whether that means it turns to 40 tables, that's fine, so long as it's relational and a query can be written.
- rudyj03 4y agoIf querying is the main goal, ingesting the pdf directly into something like elasticsearch or splunk would be much simpler. This of course doesn't meet the checkbox of being open source though.
- chaps 4y agoThis is a very... engineer... approach. Not something that I would want to ever do with the intent of exploratory analysis and eventual sharing.