4 ms·
Did you do use the same method, i.e. split by chunks each article and vectorize each chunk?
by hivacruz 2y ago
Did you do use the same method, i.e. split by chunks each article and vectorize each chunk?
- gfourfour 2y agoYes
- dudus 2y agoThat's the only way to do it. You can't index the whole thing. The challenge is chunking. There are several different algorithms to chunk content for vectorization with different pros and cons.
- minimaxir 2y agoYou can do much bigger chunks with models that support RoPE embeddings, such as nomic-embed-text-1.5 which has a 8192 context length: https://huggingface.co/nomic-ai/nomic-embed-text-v1.5 https://huggingface.co/nomic-ai/nomic-embed-text-v1.5 In theory this would be an efficiency boost but the performance math can be tricky.
- qudat 2y agoAs far as I understand it, context length degrades llm performance, so just because an llm "supports" a large context length it basically just clips a top and bottom chunk and skips over the middle bits.
- rahimnathwani 2y agoWhy would you want chunks that big for vector search? Wouldn't there be too much information in each chunk, making it harder to match a query to a concept within the chunk?
- nostrebored 2y agoThe problem is that often semantic meaning depends on state multiple paragraphs or sections away. This is a coarse way to tackle that