4 ms·
I’m surprised no one has commented on the context size limitations of these offerings when comparing to the other models. The sliding window technique really do
by ComputerGuru 3y ago
I’m surprised no one has commented on the context size limitations of these offerings when comparing to the other models. The sliding window technique really does effectively cripple its recall to approximately just 8k tokens which is just plain insufficient for a lot of tasks.
All these llama2 derivatives are only effective if you fine tune them, not just because of the parameter count as people keep harping but perhaps even more so because of the tiny context available.
A lot of my GPT3.5/4 usage involves “one offs” where it would be faster to do the thing by hand than to train/fine-tune first, made possible because of the generous context window and some amount of modest context stuffing (drives up input token costs but still a big win).
- tbalsam 3y agoA lot of attention window stuff is fluff IMPE, just gotta look at the benchmarks regardless of what the raw numbers sat.
- ComputerGuru 3y agoBenchmarks are, by definition, artificial. I'm speaking from real-world experience, not from "based off the architecture, here's my guess" and lack of context or poor recall within the "supported" context window is a real problem.
- tbalsam 3y agoYes, this is partially due, among other reasons, to the fact that inference induces a domain shift compared to training due to the use of teacher forcing.
- potatoman22 3y ago8k tokens is about 6000 words, which is more than enough for most classification tasks. Maybe it's not enough for something like story writing, but I feel like it's enough context for most business use cases.
- ComputerGuru 3y agoUsing an llm like Mistral or GPT for solely for classification is like using a jackhammer to drive framing nails. A lot of business needs require feeding the llm plenty of docs to analyze and then extract something out of, summarize, draw relationships between, etc and almost none of that can be done in 6k words. I can't even use 8k (or so-called "32k") models to analyze a moderate length Wikipedia article.