4 ms·
Every long context sucks right now. All the model providers benchmark on fact recall which is very limited. Actual ability to do anything complicated beyond 16k
by tempusalaria 3y ago
Every long context sucks right now. All the model providers benchmark on fact recall which is very limited. Actual ability to do anything complicated beyond 16k tokens is not present in any current model I have seen.
- ukuina 3y agoThis is not current. GPT-4-Turbo (128k) has lossless recall to the first 64k input tokens and produces output indistinguishable from GPT-4 (32k), though both are limited to 4k output tokens. Several downsides: Recall accuracy past the first 64k tokens suffers badly; Cost is astronomical; Response latency is too high for most interactive use-cases. I would point out the astounding leap in input context in just one year. Should we assume effectively-infinite (RAG-free) context in the near-future?
- anoncareer0212 3y agoThis is grossly untrue in a way that denotes surface-level familiarity on several fronts You're referring to the needle-in-a-haystack retrieval problem. Which the person you're replying to explicitly mentioned is the only benchmark providers are using, for good reason. Consider the "translate Moby Dick to comedic zoomer" problem. This does not even come remotely close to working unless I do it in maximum chunks of 5,000 tokens. Consider the API output limit of 4096 tokens, across all providers. And no, you shouldn't assume effectively infinite (RAG free) context in the near future. This time last year, Anthropic was demonstrating 120,000 token context. It released 200K a few weeks ago. And runtime cost scales with N^2.