4 ms·
Interesting, is this based on an external Vector DB to store and process the PDF?
by LoMoGan 1y ago
Interesting, is this based on an external Vector DB to store and process the PDF?
- mingtianzhang 1y agoThanks for the great question! We actually use a reasoning-based, vectorless approach. In short, it follows this process: 1. Generate a table of contents (ToC) for the document. 2. Read the ToC to select a relevant section. 3. Extract relevant information from the selected section. 4. If enough information has been gathered, provide the answer; otherwise, return to step 2. We believe this approach closely mimics how a human would navigate and read long PDFs.
- LoMoGan 1y agoSounds interesting, will try it out.
- mingtianzhang 1y agoThanks, any feedback is welcome!