3 ms·
How are you guys deciding what parts of a document to turn into embeddings? I've heard paragraph embeddings aren't that reliable, so I'm planning on using tf-id
by gregw134 3y ago
How are you guys deciding what parts of a document to turn into embeddings? I've heard paragraph embeddings aren't that reliable, so I'm planning on using tf-idf first to extract keywords from a document, and then just create embeddings from those keywords.
- nestorD 3y agoTake your document, cut it (cleanly!) into pieces small enough to fit into your sentence embedder's context window, and generate several embeddings that all point to the same document. I would recommend against merging (averaging, etc.) the embeddings (unless you want a blurry idea of what your document contains), as well as feeding very large pieces of text to the embedder (some models have massive context lengths, but the result is similarly vague).
- gregw134 3y ago> [recommend against] feeding very large pieces of text to the embedder Sounds right, I've heard this from multiple sources. That's why I'm leaning towards just embedding the keywords.
- Havoc 3y agoWhen trying to find similarity between whole docs one would feed the entire doc though, right?
- therealdrag0 3y agoKeyword would leave you with semantics of word definitions but lose sentence meaning/context right?
- gregw134 3y agoI'm sure it would lose a ton of meaning, but for me it's easier to fit into a traditional search pipeline.