9 ms·
This is nice, but it can get quite expensive. Let's say I have a book and I want to ask multiple questions about it. Every query will pay the price of the book
by justanotheratom 3y ago
This is nice, but it can get quite expensive.
Let's say I have a book and I want to ask multiple questions about it. Every query will pay the price of the book's text. It would be awesome if I could "index" the book once, i.e. pay for the context once, and then ask multiple questions.
- wahnfrieden 3y agoNot sure about this one but you can usually ask multiple questions in one shot at least
- minimaxir 3y agoGeneration is more expensive than the prompt input (for Claude v1, generation is 3x the cost; for GPT-4 it's 2x the cost) It makes the economics slightly trickier.
- newhouseb 3y agoI wonder why this is? Naively there's no difference between the two from a transformer standpoint. Perhaps it's because under the hood there's additional safety analysis/candidate generate that is resource intensive?
- space_fountain 3y agoWell each additional token generated requires rerunning the model right to find the next likely token given the previous one
- newhouseb 3y agoNaively, yes, but you can cache the bulk of that "rerunning" [1]. That said the (non-flash) attention costs go up with the length of the sequence so perhaps this is just a simpler way to approximate these costs. [1] https://kipp.ly/blog/transformer-inference-arithmetic/ https://kipp.ly/blog/transformer-inference-arithmetic/
- pyth0 3y agoNormally the inputs are padded out to the context length [1] and so the cost to embed 1 token or N tokens is the same. The output is produced token-by-token and so the amount of GPU time increases with the number of output tokens. [1] I'm not sure if these huge context lengths are achieved the same way (i.e. a single input vector of length N) but given the cost is constant for input I would assume the resource usage is too.
- newhouseb 3y agoThis doesn't match my mental model (or implemented model in the case of GPT2) of how self-attention works (you need to calculate the residual stream for each individual token, attending to all prior tokens before it). Have a link?
- pyth0 3y agoI work on infrastructure for serving large language models but I don't have any background in ML, so my perspective is looking at these models as a black box (and also conversations with the people that do the ML stuff). It is the case in practice at least from a latency side that with a fixed context length N, embedding any number of tokens from 0 to N takes the same amount of time. Perhaps it's a difference between the conceptual and actual implementation on GPU? edit - This occurred to me after the fact but I wonder if the difference is that the use case I work with is processing batches of many different embedding requests (but computed in one batch), therefore it has to process `min(longest embedding, N)` tokens so any individual request in theory has no difference. This would also be the case for Anthropic however.
- newhouseb 3y agoAh, you're thinking about embeddings which are basically the encoder stack on a traditional transformer architecture. Modern GPT-like models (including Claude), however, drop the encoder and use decoder-only architectures. I could imagine something where encoders pad up to the context length because causal masking doesn't apply and the self attention has learned to look across the whole context-window.
- atq2119 3y agoIt's because the input tokens can be batch-processed in a single forward pass through the model, while generating tokens requires one forward pass through the model per token. If you do the math of how much memory bandwidth is required by a forward pass vs. how much compute, you'll see that inference is entirely limited by memory bandwidth and will use compute resources very inefficiently. In contrast, input processing is able to fully use the available compute. Of course, there are ways to mitigate this problem, like processing multiple token streams in parallel, but the fundamental problem remains.
- newhouseb 3y agoAh, this makes total sense. I was thinking about FLOPs in the abstract and not about the wall-clock time. Thanks for the explanation.
- tikkun 3y agoWith embeddings, you essentially can. Group the book into sections, embed each section, then when you do a prompt, add in the N most similar embedded sections to your prompt.
- adamgordonbell 3y agoWhat if the question is "What are the main themes of this work?" Or anything where the question answer isn't 'close' to the words used in the question? How well does this work vs giving it the whole thing as a prompt? I assume worse but I'm not sure how this approach compares to giving it the full thing in the prompt or splitting it into N sections and running on each and then summarizing.
- jtlicardo 3y agoYou pretty much summed up the drawbacks of the embeddings approach. In my experience it's pretty hard to extract the relevant parts of text, especially when the text is uniform.
- abraxas 3y agoYou could do multi level summaries etc but yeah this is all just band aids around token limits.
- Spivak 3y agoI don't think it's as much of a band-aid as it first appears since this roughly mimics how a human would do it. The problem is that humans have continuous information retrieval and storage where the current crop of embedding systems are static and mostly one shot.
- crucialfelix 3y agoHumans have limited working memory, they quickly forget short term memory (unless it's super significant) and our long term memory fades selectively if not reactivated or significant (intense). This weird leaky memory has advantages and disadvantages. Forgetting is useful, it removes garbage. Machine models could vary the balance of temporal types, drop out Etc. We may get some weird behavior. I would guess we will see many innovations in how memory is stored in systems like these.
- fdgsdfogijq 3y agoThe price on this will plummet over the next few years, the economic benefits are too large
- nr2x 3y agoYes, but physics trumps economics.
- moffkalast 3y agoThe economic benefits of mining asteroids are also too large to ignore yet here we are, levelling villages to dig for coal. Just a few manufacturers hold the effective cartel monopoly on LLM acceleration and you best bet they will charge out the ass for it.
- skybrian 3y agoI'm wondering what level you're thinking. Cloud vendors? GPU vendors? Fabs?
- moffkalast 3y agoGiven what's used right now to my knowledge, the main ones would be Nvidia's tensor cores, Apple's M chips and Google's cloud TPUs. All of that's TSMC I think?
- modernpink 3y agoMarket competition and innovation in both ML and hardware has consistently driven down the price of AI in the past decade. You only have to look at where we are with capabilities today compared to ten years ago when CIFAR100 classifiers were the state of the art. Barring a Chinese invasion of Taiwan, these APIs will halve in price over the next year.
- deleted 3y ago[deleted]
- moffkalast 3y ago
- pyth0 3y agoThis more or less is already a thing and it's called RAG [1][2]. It essentially allows you to have a database of embeddings (in this case your book) from which a model can pull knowledge from while producing answers. As for the standard operation of these generative models, the context window is the only working memory it has and so it must see the entire text each time. [1] https://arxiv.org/abs/2005.11401 https://arxiv.org/abs/2005.11401 [2] https://huggingface.co/docs/transformers/model_doc/rag https://huggingface.co/docs/transformers/model_doc/rag
- m1sta_ 3y agoCam you help me understand this? The research appears to be from a few years ago. Can this be used with Claude (for example)? How is it different to the approach many people are taking with vector stores and embeddings?
- make3 3y agoit's not different. RAG is a way to train embedding stores end to end
- make3 3y agosomehow got down voted on something I'm a professional expert at
- pyth0 3y agoOther people seem to be suggesting that the user would do the retrieval of the relevant parts of the book from a vectordb first, and then feed those sections along with the question as the prompt. Conceptually it is very similar (and it too uses vector database), but with RAG it would happen as part of the inferencing pipeline and therefore achieve better performance than the end user emulating it.
- ukuina 3y agoYep, but your retrieval from the vector DB becomes your relevancy bottleneck.
- mikrl 3y agoThe analogy I can think of here is a pointer, but AFAIK the context would always need to go along with the prompt unless you could tweak internal state to bias towards the context. Otherwise, it might make sense to have a separate routine which compresses the context as efficiently as possible. Auto encoder?
- make3 3y agoYes, caching the states of the sequence would make sense. An issue is that it's still more expensive to compute the new tokens even if you cache the states viewed so far