4 ms·
I'm not sure that I understand what we're parsing to. Like on the website, I see supported types, but that looks like the parsable types, no? What kind of struc
by bpev 2y ago
I'm not sure that I understand what we're parsing to. Like on the website, I see supported types, but that looks like the parsable types, no? What kind of structured representation is outputted? And can we guide what that structure looks like?
- adithya-s-k 2y agoYes, the current implementation of the repository converts any data primarily into strctured markdown text. The next stage will involve prompt guides or schema-guided structure extraction. Let's say you are processing a lot of research PDFs and want to convert them into clean markdown that best represents the content. Now, let's say you want to extract the authors, abstracts, captions, and store images. The extraction engine we are currently working on will help you with that.
- xigoi 2y ago“structured Markdown” sounds like an oxymoron.