3 ms·
Could you share more about your strategy or the approach you took for chunking the articles? I'm curious about the criteria or methods you used to decide the bo
by netdur 3y ago
Could you share more about your strategy or the approach you took for chunking the articles? I'm curious about the criteria or methods you used to decide the boundaries of each chunk and how you ensured the chunks remained meaningful for the search functionality. Thanks!
- mattkevan 3y agoAs this was my first attempt, I decided to take a pretty basic approach, see what the results were like and optimise it later. Content is stored in Django as posts, so I wrote a custom document reader that created a new LlamaIndex document for each post, attaching the post id, title, link and published date as metadata. This gave better results than just loading in all the content as a text or CSV file, which I tried first. I did try with a bunch of different techniques to split the chunks, including by sentence count and a larger and smaller number of tokens. In the end I decided to leave it to the LlamaIndex default just to get it working.
- drittich 3y agoI wrote a C# library to do this, which is similar to other chunking approaches that are common, like the way langchain does it: https://github.com/drittich/SemanticSlicer https://github.com/drittich/SemanticSlicer Given a list of separators (regexes), it goes through them in order and keeps splitting the text by them until the chunk fits within the desired size. By putting the higher level separators first (e.g., for HTML split by <h1> before <h2>), it's a pretty good proxy for maintaining context. Which chunk size you decide on largely depends on your data, so I typically eyeball a sample of the results to determine if the splitting is satisfactory. You can see the separators here: https://github.com/drittich/SemanticSlicer/blob/main/SemanticSlicer/Separators.cs https://github.com/drittich/SemanticSlicer/blob/main/Semanti...