3 ms·
Can't someone expand on this > Chunking is more or less a fixable problem with some clever techniques: these are pretty well documented around the internet; C
by ratedgene 2y ago
Can't someone expand on this
> Chunking is more or less a fixable problem with some clever techniques: these are pretty well documented around the internet;
Curious about what chunking solutions are out there for different sets of data/problems
- pphysch 2y agoMost data has semantic boundaries: whether tokens, words, lines, paragraphs, blocks, sections, articles, chapters, versions, etc. and ideally the chunking algorithm will align with those boundaries in the actual data. But there is a lot of variety.
- hansvm 2y agoIt's only "solved" if you're okay with a 50-90% retrieval rate or have particularly nice data. There's a lot of stuff like "referencing the techniques from Chapter 2 we do <blah>" in the wild, and any chunking solution is unlikely to correctly answer queries involving both Chapter 2 and <blah>, at least not without significant false positive rates. That said, the chunking people are doing is worse than the SOTA. The core thing you want to do is understand your data well enough to ensure that any question, as best as possible, has relevant data within a single chunk. Details vary (maybe the details are what you're asking for?).
- haolez 2y agoI had some success with simple aliasing at the beginning and end of chunks. In my next project, I'll try an idea that I saw somewhere: 1. do naive chunking like before 2. calculate the embeddings of each chunk 3. clusterize the chunks by their embeddings to see which chunks actually bring new information to the corpus 4. summarize similar chunks into smaller chunks Sounds like a smart way of using embeddings to reduce the amount of context misses. I'm not sure it works well, though :)