3 ms·
Digest, an AI document parser with referencing
- adriankibet 3y ago[dead]
- poppingtonic 3y agoHello HN, co-founder (Brian Muhia), here to answer any questions!
- Joseph_Gicuguma 3y agoCongratulations on the launch of Digest, Adrian! The ability to extract specific information from documents by simply asking questions is impressive. I'm curious about the scalability of your product. How does Digest handle large document archives, and what measures are in place to ensure efficient search and retrieval of data from extensive databases?
- adriankibet 3y agoHey, thanks Joseph, appreciate it. So in that particular sense it’s very scalable/efficient, the LLM doesn’t really need to scale its parsing at all because of the preprocessing we do beforehand. When you upload your doc we store it as paragraphs and keep the actual doc in cold cloud storage. Just below the question bar a question in the archive view, there’s a selection slider for the number of paragraphs you want used as context to answer the question, this can be adjusted upwards if you think more context would help but 40 paragraphs is the number we settled on after testing because lower than that and you start getting inaccuracies. Once you ask a question Digest checks which 40 or more paragraphs are likely to answer the question then sends those as the context for the LLM to answer with so regardless of the size of the archive the speed of the search is fairly consistent because its always ~40 paragraphs (with some adjustments for paragraph size). Thanks for the question and I hope that helps.