5 ms·
Agentic Context Management: Memory and Cost as Architecture Problems
- gdad 1mo agoI also wrote a shorter preview here: https://www.maximem.ai/blog/agentic-context-management-paper https://www.maximem.ai/blog/agentic-context-management-paper
- respectattentio 1mo agoI like to start with memory engineering then reach full system then reducing costs. This allows unlocking full potential of agents.
- gdad 1mo agoInteresting. Where can I read more about this?
- respectattentio 1mo agoI came up with this after building a few systems. For me, if I put costs in my architecture, I'm limited heavily and that can easily change how memory is shaped dramatically. The opposite, putting memory in architecture, is not true. Memory engineering first, then full scale in the system then costs considerations. In addition, I believe this is future friendly. Because AI is advancing and getting smarter and cheaper everyday. I couldn't find a guide on this so I share my basic thoughts.
- samyakk 1mo agoACM, that's the term that I'd been looking for - and your paper explains it clearly. At the end, most of LLM problems are context problems. Getting the correct knowledge into its context window without overpopulating it is the actual engineering effort for most agents. And the solution you present seems promising. Both compaction with validation and predictive fetching are the way to go. I do not want to write an implementation for this myself, and if Synap is that implementation, I'd like to ask you a few questions: 1. Does it work with context that's not just agent conversations, but rather documents? 2. Is it better than RAG on large dataset? 3. What does on-prem options look like?
- yeasin-arafat 1mo ago[flagged]
- gdad 1mo agoThanks Samyakk! 1. Yes, works on docs, agent conversations, human-conversations from different sources (Slack, JIRA, etc.). We have connectors for some of these as well; so it is plug and play 2. conventional RAG recall accuracy is quite low (50-60%) and latency is pretty high (seconds). But worst is the precision; you end up context stuffing to get acceptable recall 3. We do offer on-prem deployments, but only on sizeable annual contracts
- tomveber 1mo ago[dead]
- trsthales 1mo ago[flagged]
- melembre 1mo agoContext drift on retries is easily the most annoying part of this setup. Locking down the tool payload schema first was the only thing that worked for us
- nullbio 1mo agoContext pollution and rot are probably more important than memory, because facts can usually be retrieved if the agent is good at following breadcrumbs. What's also the biggest killer is code rot. Agents are particularly good at death by thousand cuts. They implement something poorly, or incorrectly, or introduce a bad pattern into the project. Then they continue to amplify that badness over time, as they continue to copy from it on subsequent work. It spreads like a virus. Keeping these seeds out of the project is very difficult, and cleaning up the rot is very difficult. It also seems like a hard problem to solve because following the existing codebase is something that is good when the code is good, but bad when it is bad. So, seemingly, the solution means more thinking and evaluation for every change that is being made.
- gdad 1mo agoTruly. Doing this for coding agents is an interesting and different shaped problem.
- ShinyLeftPad 1mo ago> They implement something poorly, or incorrectly, or introduce a bad pattern into the project. Then they continue to amplify that badness over time, as they continue to copy from it on subsequent work. It spreads like a virus. > Keeping these seeds out of the project is very difficult, and cleaning up the rot is very difficult. My "aha" moment was when I realized this goes for all spheres of life where this tech is/will be introduced.
- taneq 1mo agoIt goes for all spheres of life, full stop. I’m not sure if agents struggle with this because they learned it from humans, or if they struggle with it because it’s a universally challenging problem, but it’s something we share with them.
- ShinyLeftPad 1mo agoThe comment highlighted how LLMs exacerbate the issue by entrenching the preexisting issues.
- miranaproarrow 1mo agoIve never read a paper cover to cover before but after wrestling with opus 5s english this paper is such a relief to read, its like my eyes has been washed off opus stink
- altwebsoftware 1mo ago[flagged]
- gdorsi 1mo agoNice, is there any harness that implements this approach?
- ShadowandPath 1mo ago[flagged]
- janeabegail1 1mo ago[flagged]
- profitplay 1mo ago[flagged]
- felixlu2026 1mo ago[dead]
- jkwang 1mo ago[flagged]
- sangwook 1mo agoIm wondering how silent information loss is detected later and what exactly the validation score measures.
- benzguo 1mo agoI've found that a simple markdown knowledgebase (with some useful extensions like semantic search & git context) is all I need to improve the memory of my agents. Even my non-coding agents have a memory repo. Here's my implementation: https://hraness.com/kb https://hraness.com/kb
- datadrivenangel 1mo ago[dead]
- alankritxghoshx 1mo ago[flagged]
- wajdym1960 1mo ago[flagged]