3 ms·
I’m surprised no one has commented on the context size limitations of these offerings when comparing to the other models. The sliding window technique really do
by ComputerGuru 3y ago
I’m surprised no one has commented on the context size limitations of these offerings when comparing to the other models. The sliding window technique really does effectively cripple its recall to approximately just 8k tokens which is just plain insufficient for a lot of tasks.
All these llama2 derivatives are only effective if you fine tune them, not just because of the parameter count as people keep harping but perhaps even more so because of the tiny context available.
A lot of my GPT3.5/4 usage involves “one offs” where it would be faster to do the thing by hand than to train/fine-tune first, made possible because of the generous context window and some amount of modest context stuffing (drives up input token costs but still a big win).
- logicchains 3y ago>The sliding window technique really does effectively cripple its recall to approximately just 8k tokens which is just plain insufficient for a lot of tasks. What are you basing this observation on; personal experience, or is there a benchmark somewhere confirming it?
- ComputerGuru 3y agoReal world testing and experience. If you only need the llm to retain the "gist" of the input tokens in order to return a "related" answer, the sliding window design is fine. But if you need actual technical analysis or tasks that involve verbatim referencing, quoting, recomposing, etc based off parts of the input documents, it doesn't work. I tried using it for "business document" use cases but have ran into this with code as well; the latter might be a better explanation given where we're having this discussion. If you only need the llm to retain the general shape of your inputs so it can reuse them to influence the output, the sliding context is fine. But if you need it to actually reuse code verbatim from the input that you fed it (or to remember the api calls and their surrounding context verbatim to recall from a sample of just one that this api must be called before that api, when the prompt includes instructions to that effect) the "decomposition" of the input tokens with the sliding model is insufficient and the llm completely fails at the assigned task.
- cameroncairns 3y agoI found this discussion on the local llama subreddit that digs a little bit more into what effects the sliding window might have, in case you or anyone else reading this comment thread finds it interesting: https://old.reddit.com/r/LocalLLaMA/comments/17k2mwq/i_dont_understand_mistral_and_context_size/ https://old.reddit.com/r/LocalLLaMA/comments/17k2mwq/i_dont_.... It refers to the original Mistral 7B though not the new Mixtral fwiw
- rockinghigh 3y agoMixtral 8X7B has a 32k-token context window. GPT-3.5 models have a context window of 4k-16k. As for the sliding window attention, the model does not lose all the information about the tokens before the sliding window. Hidden states store information about past tokens.
- ComputerGuru 3y agoThe sliding window does not lose the data but it does "decompose" it so that it can't be recalled verbatim. For analyzing code (feed it n classes and ask it to create a class using all of them to accomplish a task) that isn't good enough. It's also not good enough for some of the "corporate business" use cases we tried putting Mistral and other sliding window models to use on, where you need it to re-use verbatim or reference specific portions of one of n documents fed into it as input tokens. Again, sufficient training can overcome these limitations. But that's only for cases where the corpus of input documents is static or at least contains significant reuse.