3 ms·
I’m struggling to understand what’s novel here. LLM-based summarization of chat history memory is a well-established technique implemented by many LLM framework
by roseway4 3y ago
I’m struggling to understand what’s novel here. LLM-based summarization of chat history memory is a well-established technique implemented by many LLM frameworks. Summarizing on every message is, as proposed in the paper, a major performance bottleneck and adds significant latency to the chat loop.
Many implementations utilize a fixed sized buffer, progressively summarizing batches of older memories when they fall out of the buffer. Ideally, this is also done out of band to the chat loop.
I’m an author of Zep[0], an open source long-term memory store, and this is how we implemented summarization.
0: https://github.com/getzep/zep https://github.com/getzep/zep
- anotherpaulg 3y agoAider does this too, using a background thread to summarize the messages older than the last N. https://github.com/paul-gauthier/aider/blob/main/aider/history.py https://github.com/paul-gauthier/aider/blob/main/aider/histo...
- px43 3y agoYeah, I'm pretty novice, but I took that hourish long Andrew Ng class on LangChain, and they covered recursive summarization as a standard memory management technique. https://www.deeplearning.ai/short-courses/langchain-for-llm-application-development/ https://www.deeplearning.ai/short-courses/langchain-for-llm-...
- deleted 3y ago[deleted]
- chrgy 3y agoExactly, There is nothing Novel here, even a middle school Chatgpt user would have known this.
- deleted 3y ago[deleted]