2 ms·
The sliding window does not lose the data but it does "decompose" it so that it can't be recalled verbatim. For analyzing code (feed it n classes and ask it to
by ComputerGuru 3y ago
The sliding window does not lose the data but it does "decompose" it so that it can't be recalled verbatim. For analyzing code (feed it n classes and ask it to create a class using all of them to accomplish a task) that isn't good enough. It's also not good enough for some of the "corporate business" use cases we tried putting Mistral and other sliding window models to use on, where you need it to re-use verbatim or reference specific portions of one of n documents fed into it as input tokens.
Again, sufficient training can overcome these limitations. But that's only for cases where the corpus of input documents is static or at least contains significant reuse.