3 ms·
When you are dealing with documents with different structures, how to do the document chunking efficiently without losing important metadata?
by zergnick 2y ago
When you are dealing with documents with different structures, how to do the document chunking efficiently without losing important metadata?
- tylersuard 2y agoFirst of all, great question. Second, we use a search service, and vectors are treated as supplementary to the text search, so chunking doesn't matter as much. We will usually take an entire PDF page and embed that, no matter what structure the data on that page is. We do keep track of the name of the document and the page number. For SQL records, we just turn each record into a text string and embed that.
- zergnick 2y agoThanks for your feedback! Could you share a bit about your team? I’m curious how many people are involved and what kinds of skills or roles are needed to make this happen.