3 ms·
I find that ingesting and chunking PDF textbooks automatically creates more of a fuzzy keyword index than a high level conceptual knowledge base. Manually curat
by deckar01 3y ago
I find that ingesting and chunking PDF textbooks automatically creates more of a fuzzy keyword index than a high level conceptual knowledge base. Manually curating the text into chunks and annotating high level context is an improvement, but it seems like chunks should be stored as a dependency tree so that, regardless of delineation, on retrieval the full context is recovered.