4 ms·
That per exchange context cost is what really puts me off using cloud LLM for anything serious. I know batching and everything is needed in the data center, and
by mcbuilder 2y ago
That per exchange context cost is what really puts me off using cloud LLM for anything serious. I know batching and everything is needed in the data center, and important for keeping around KVQ cache, you basically need to fully take over machine to get an interactive session to get the context costs to scale with sequence length. So it's useful, but more in the case of a local LLaMA type situation if you want a conversation.
- falcor84 2y agoI wonder if we could implement the equivalent of a JIT compilation, whereby context sequences which get repeatedly reused would be used for an online fine-tuning.
- HumanOstrich 2y ago[flagged]
- nostrebored 2y agoThey are asking if you can take the context being passed per interaction and train it into a session in real time (via an online algorithm). Essentially bake the context passed in to the attention layer so that you can pass only the relevant chat context. Your post wasn’t a particularly charitable interpretation.
- sp332 2y agoNo, but you can just cache the state after processing the prompt. https://github.com/ggerganov/llama.cpp/tree/master/examples/main#prompt-caching https://github.com/ggerganov/llama.cpp/tree/master/examples/...
- neverokay 2y agoIt makes building any app that requires generous user prompting impossible to build for regular developers (cloud pricing). $20 hosting can serve thousands of users per month. $20 llm sub services just one person. This is fucking impossible.
- cyrusalwati 2y agoHow are you hosting for $20 per month? I've been under the impression that I'd need to pay closer to $200 to self host. I'm building a small assistant tool and thought I'm forced to use APIs!