6 ms·
i'm curious about the chunk splitting approach you mentioned. how do you determine the optimal chunk size for processing? seems like there could be a tradeoff b
by ramonverse 2y ago
i'm curious about the chunk splitting approach you mentioned. how do you determine the optimal chunk size for processing? seems like there could be a tradeoff between context preservation and processing efficiency. have you experimented with different chunk sizes and their impact on the quality of the final output? this could be really important for handling things like long-range dependencies in the text.
- eigenvalue 2y agoI just tried messing around with the chunk size until the results looked best to my eye. One important thing to note is that each chunk includes a portion of the previous chunk as context. But yes, it’s not going to have context that spans the entire document. For that you would need a final step that takes the entire output. One thing I learned from this is that you can’t ask too much at once from these value tier models. If you keep your requests modest and focused, they do really well. When you try to cram too much complexity and too many rules/requests at once, they start messing up and leaving stuff out and hallucinating.