5 ms·
The longer the context the more backtracking it needs to do. It gets exponentially more expensive. You can increase it a little, but not enough to solve the pro
by dudus 2y ago
The longer the context the more backtracking it needs to do. It gets exponentially more expensive. You can increase it a little, but not enough to solve the problem.
Instead you need to chunk your data and store it in a vector database so you can do semantic search and include only the bits that are most relevant in the context.
LLM is a cool tool. You need to build around it. OpenAI should start shipping these other components so people can build their solutions and make their money selling shovels.
Instead they want end user to pay them to use the LLM without any custom tooling around. I don't think that's a winning strategy.
- tom1337 2y ago> Instead you need to chunk your data and store it in a vector database so you can do semantic search and include only the bits that are most relevant in the context. Isn't that kind of what Anthropic is offering with projects? Where you can upload information and PDF files and stuff which are then always available in the chat?
- cma 2y agoThey put all the project in the context, works much better than RAG when it fits. 200k context for their pro plan, and 500K for enterprise.
- gcr 2y agoThis isn't true. Transformer architectures generally take quadratic time wrt sequence length, not exponential. Architectural innovations like flash attention also mitigate this somewhat. Backtracking isn't involved, transformers are feedforward. Google advertises support for 128k tokens, with 2M-token sequences available to folks who pay the big bucks: https://blog.google/technology/ai/google-gemini-next-generation-model-february-2024/ https://blog.google/technology/ai/google-gemini-next-generat...
- dartos 2y agoDuring inference time, yes, but training time does scale exponentially as backpropagation still has to happen. You can’t use fancy flash attention tricks either.
- thunderbird120 2y agoNo, additional context does not cause exponential slowdowns and you absolutely can use FlashAttention tricks during training, I'm doing it right now. Transformers are not RNNs, they are not unrolled across timesteps, the backpropagation path for a 1,000,000 context LLM is not any longer than a 100 context LLM of the same size. The only thing which is larger is the self attention calculation which is quadratic wrt compute and linear wrt memory if you use FlashAttention or similar fused self attention calculations. These calculations can be further parallelized using tricks like ring attention to distribute very large attention calculations over many nodes. This is how google trained their 10M context version of Gemini.
- dartos 2y agoI may be missing something, but I thought that each context token would result in an 3 additional parameters per context token for self attention to build its map, since each attention must calculate a value considering all existing context
- upghost 2y agoSo why are the context windows so "small", then? It would seem that if the cost was not so great, then having a larger context window would give an advantage over the competition.
- thunderbird120 2y agoThe cost for both training and inference is vaguely quadratic while, for the vast majority of users, the marginal utility of additional context is sharply diminishing. For 99% of ChatGPT users something like 8192 tokens, or about 20 pages of context would be plenty. Companies have to balance the cost of training and serving models. Google did train an uber long context version of Gemini but since Gemini itself fundamentally was not better than GPT-4 or Claude this didn't really matter much, since so few people actually benefited from such a niche advantage it didn't really shift the playing field in their favor.
- Melatonic 2y agoSeems like a good candidate for a "dumb" AI you can run locally to grab data you need and filter it down before giving to OpenAI
- solarkraft 2y ago> you need to chunk your data and store it in a vector database so you can do semantic search and include only the bits that are most relevant in the context Be aware that this tends to give bad results. Once RAG is involved you essentially only do slightly better than a traditional search, a lot of nuance gets lost.
- rahimnathwani 2y agoThis depends on the amount of context you provide, and the quality of your retrieval step.
- hackernewds 2y agoI don't know whether using exponential in the general English language usage of the word, but it does not get exponentially more expensive