3 ms·
Important to note though, that even if they say it has 200k context, there's a good chance it won't be able to use all of it. For example, here is an experimen
by Jackson__ 3y ago
Important to note though, that even if they say it has 200k context, there's a good chance it won't be able to use all of it.
For example, here is an experiment on long context recall for Claude and GPT4:
https://twitter.com/GregKamradt/status/1727018183608193393 https://twitter.com/GregKamradt/status/1727018183608193393
https://twitter.com/GregKamradt/status/1722386725635580292 https://twitter.com/GregKamradt/status/1722386725635580292
Even though claude claims 200k tokens context, perfect memory stops at about 19k tokens.
I'd personally expect the Yi-34b models to perform even worse at the task, though I have not tested it.
- xena 3y agoYeah, I'm planning on testing it as soon as I figure out how to properly prompt engineer an untuned foundation model. GPT-4 Turbo 128k also seems to lose coherence around 20k tokens based on my testing so far.
- wahnfrieden 3y agohave you tested?
- taf2 3y agoYeah we tried Claude 2 to summarize multiple conversations and if we give it the instructions at the start and then show it the conversations it almost always forgot the original instruction and ends up outputting something related to the conversation so long context whole nice seem to break down very quickly. What I found worked better was to provide the instructions at the end but given the in ability to follow the starting instructions my guess is much of the early conversations are lost… so yeah not so sure big context is a silver bullet