7 ms·
Agent memory as a file format
- tylerjharden 1mo agoWouldn't Avro or Parquet be solid choices for something like this, or am I out of date & out of touch?
- laruss5 1mo ago[dead]
- pietz 1mo agoI'm not convinced an unstructured collection of memory files is the way to go at all.
- dist-epoch 1mo agoIf you look at how agents navigate source code, they do not look at directory names, and drill down into the ones with plausible names, instead the grep the whole repo for plausible keywords. Of course, ideally your data would be structured, but the agents will mostly be grepping anyway, and maybe look at sibbling files.
- DarmokTanagra 1mo agoits the same way they crawl websites, its horribly inefficient
- deleted 1mo ago[deleted]
- JustFinishedBSG 1mo agoThat's a whole lot of text to say "it's markdown".
- thm 1mo agoA whole lot of text to say "Just use text".
- andyfilms1 1mo agoIt feels like 50% of AI "progress" is just finding different ways to say "tell the model things in English."
- cyanydeez 1mo agoultimately, it's all a bunch of text; and like playdo, whats interesting is all the different types of holes you can squeeze it through. But it's still text. You show people how the sausage is made and either they're confused or they're horrified.
- wasabi991011 1mo agoThat's only half. It's a while lot of text to say "it's markdown + RAG/semantic search". The markdown explanation I was also unimpressed with, but RAG over hyperlinks is convincing to me.
- pavo-etc 1mo agoI've come to a similar lofi solution for my agent fleet. Markdown wiki with simple querying is decently effective as a memory system. Setting up a skill that can effectively reduce a session into useful long term lessons is the easiest unlock for these systems.
- ltsSmitty 1mo agoThis was a compelling writeup to me. I read through the spec and found it easy to understand and make sense of. I wonder how much my system needs something like this. Between the invisible system memory of my random chats with Gippity, my Matt Pocock skills saving terminology and plans, and whatever else Cursor and Codex do, I don't think I feel a need for more agent memory. I do like how it's exposed and searchable, and not invisible. But I honestly just send my questions/tasks away to my magic agent and eventually it gets it right anyway; do I need more discrete memory my team has to maintain? (That's an earnest question, not disregard for this)
- DarmokTanagra 1mo agothe new OpenAI spec is agent memory as file names
- nullbio 1mo agoWhich spec are you referring to?
- DarmokTanagra 1mo agothe IM1 HuggingFace spec
- titzer 1mo agoAgent memory is to computer memory is what Mongo DB is to relational database. Incredible to watch things come full circle. Next thing you know, someone is going to figure out a binary encoding.
- dominotw 1mo agoagents can call relational databases fine. infact, really good at sql if you can store your memories into that format.
- DarmokTanagra 1mo agothat sounds horribly token inefficient, just create a tool call if you are all in on the agentic approach and hide memory retrieval behind an optimized api
- phoghed 1mo agoThey already have. The tool is called bash and optimized api is grep, the storage is the file system. We’ve been here many months. What people are exploring are other options as far as I can tell. What you are offering is “don’t do that, this already works”. Which I guess is fine, but apparently not everyone is fully satisfied with the current generation of tooling.
- DarmokTanagra 1mo agoI was responding specifically to the parent comment that was suggesting generating ad-hoc sql. Also fwiw grep is pretty poorly suited to semantic search and will only return the most basic of matches. If you are really trying to build a useful memory search tool there are much better options than plain text search.
- phoghed 1mo agoYeah but probably the whole of the memories for any given project will likely fit into context, and the agent can’t extract whatever it wants for the given task. Grep only comes in if they choose to narrow down the candidate memory files, and they are usually searching many keywords and synonyms. Of course a real search system with ranking and whatnot would be better, but you’re paying a different cost there. I’m personally more interested in how the memory files get created and updated and generally managed, retrieving the correct ones doesn’t seem to be a big problem at the moment. I primarily use gh copilot and they are by default hidden from the user. It’s also not great that they aren’t in the repo and every contributor has their own set of memories of different freshness, likely conflicting.
- guhcampos 1mo agoWhat the author suggests is remarkably close to the proposition of OpenViking. I've been testing a few memory solutions and OpenViking is one of my favorites so far.
- strix_varius 1mo agoI'm looking in the same space: have you found anything else you're considering beyond OpenViking?
- MichaelGlass 1mo agoI see a lot of claims in this article without ... any proof? Both can be true: - It's useful to anthropomorphize agents when predicting behavior and - we have to use specific language to specify what we mean. What does the author mean by "confuse the models" ? Are they talking about not picking right information? Picking the wrong information? Losing their previous context / task? Part of setting up a proper eval is also deciding what we actually mean ourself. What are we actually optimizing for? It's not, e.g. % confusion, %rubbish, etc. The article does point to it: retrieval latency, accuracy, etc.
- bensyverson 1mo agoIt’s good that a lot of people are trying a lot of things when it comes to agentic memory. Sadly none of it represents a complete solution at this time. But we need the experimentation.
- dominotw 1mo agoI still dont understand what "agentic memory" is . agents can already call sql / rag and grep through files or whatever. why is "agentic memory" a special thing.
- elliotbnvl 1mo agoYou could think of it as a more efficient indexing format for data the agent has access to. It’s badly needed as a better representation or pointer map would improve recall time and comprehensiveness dramatically without having to invest further in model training.
- TudorAndrei 1mo agothey just want to reinvent information retrieval from first principles; also the fact that everyone forgets to check what currently exists and reinvents hexagonal wheels for the sake of agentic development
- phoghed 1mo agoRight now you have an opportunity to enlighten everyone rather than talk down to them.
- calpaterson 1mo agoJust retrieving usually doesn't count as memory. "Memory" tends to imply writing too. And I agree: it's not a very special thing. That's why I propose: Markdown + a simple embedding.
- edgyquant 1mo agoMemory isn’t the same thing as rag it’s usually just a text based index of past events and the llm reads it and decides what’s important rather than querying a db
- Kuyawa 1mo ago[flagged]
- skapadia 1mo agoMemory is not just a matter of retrieval, it's also a matter of knowing what to retrieve and when.
- docheinestages 1mo ago> How can I judge what is a good memory to store? How can I avoid filling my memory with crap? > This is a common fear with memory systems but doesn't really apply to memoryfields. Irrelevant material is simply never surfaced by the semantic search. This is so wrong. The Achilles' heel of this approach is the RAG. What makes it worse is having lots of memories that are outdated, wrong, hallucinated, or irrelevant. Nothing beats curated data. Memory should be regularly reviewed, compacted, and cleaned up if it's no longer valid.
- artyomsv 1mo agoCheapest version of that review is already sitting in the format. Frontmatter carries created and updated, so you can sort by staleness and drop whatever nothing touched in months
- dfrancislyondfl 26d ago[flagged]
- dataviz1000 1mo agoDoes anyone else not use memory? I find once there is one poisoned line of text it negatively affects everything else downstream. Instead, I use a temp/ folder with documents and use different files for different agents and models. Then I have to constantly prune and delete the files. Any information that can be extrapolated is just noise which negatively affects the agent. If you have a definition of a database structure and it has been implemented, that information should not be contained in any text document -- it is noise, will drift, and be impossible to debug why the agent keeps producing undesired behavior. I have a ~/Projects folder. For example, I use Playwright with Chrome DevTools Protocol in order to do performance testing and leak detection. There is a script that handles this. My prompt is "Search ~/Projects for perf testing with CDP and Playwright and implement here". Point being, if I need anything I point to a resource or ask to search a resource and it will find it quick and, most importantly, tends to improve it every iteration. If I was in an institution, I would have a repository and would rather just point the resource and say use that than have memory of it locally.
- jjice 1mo agoI keep it on for my web chats, but I think I dislike it more than it helps. I'll ask a question about something and it'll find a way to tie it back to something from three months ago thinking it was a deep project I was working on, instead of what it really was: an inane question I was curious about. I turn it off for local agents because I bounce around a few and I really don't like the mostly implicit nature of it. I want to write my instructions in version control if I have anything to say consistently to an agent.
- calpaterson 1mo agoWhat I propose is only a very slightly more formal version of what you describe. Just to start with: memoryfields are possible to use in a server/client system. That was a key aim and I do already use them over Amazon S3 (though not always). I started, like you did, with a personal library of prompts. But the issue is that as your library of little pieces of prompts increases a) you get tired of constantly editing them yourself b) you have no easy way to export and share them with others c) it's frustrating that the agent doesn't "automatically" find your little bit of prompt on X even when clearly it is relevant - hence sem search. I think a lot of people are still using the "personal library of bits of prompt" model. It is ok. But I wanted to propose an minimal, interchangeable standard for sharing them. So the idea of being an institution and having a shared memoryfield: that's something I want as well! The spec, feedback greatly welcome: https://github.com/calpaterson/memoryfield-spec/blob/main/SPEC.md https://github.com/calpaterson/memoryfield-spec/blob/main/SP...
- CGamesPlay 1mo agoAre embeddings useful for something of the scale compared to just keyword search (aka grep)?
- calpaterson 1mo agoI find that semantic search is substantially better than keyword search even for small corpuses. Being able to find "related" material that doesn't match the keyword is a big advance over traditional full text search.
- morelandjs 1mo agoWas ready to write something snarky because this is essentially RAG, but I think the author is getting at some subtle details which are seemingly important. - memory systems are a specific type of knowledge base where you generate all the documents. You might as well generate them to be less than your embedding token limit to obviate the need for chunking. - embedding models are getting better and are no longer just semantic averaging. - small models are getting dirt cheap, making parallel reads cost manageable What they describe is sort of the simplest architecture that takes advantage of these observations. I believe them when they say it works well. I do suspect though that things like keyword lookup will completely fail if every memory is just a vector. Hence why something like Typesense hybrid search can still be useful.
- nzach 1mo agoI'm starting to think that 'memory' may be the wrong analogy for what we want. I do think that having a set of token that are highly personalized to your project and to way you work is beneficial. I also think that the idea that this set of token will be constructed in the background without any work from the user is really appealing. So it's understandable that the 'memory' analogy became so popular. But in my experience having a really good AGENTS.md file almost always produce better results than enabling memory. Maybe we should start to think about how we 'train'/'onboard' agents into our projects, in a similar way that we do for new co-workers. Imagine if we could send the agent to our repo and ask it to learn our patterns and in the end we could quiz the agent to gauge how much it actually understood the project. Once he 'understands' the project we can start to use it to help with development. In a very small scale (example, individual new features) I will sometimes ask the agent to explain me how things work (even though I already know how it works) so I can 'prime' the agent context with good data before starting any real work. But I'm not sure if this approach could be reliably scaled to work with any repo for any kind of work.
- DanielHB 1mo agoWhen I was doing a really large refactor across the codebase I told Claude Code to explain how certain things worked currently and how I wanted things to look like after the migration and a plan on how to get there. Then I did several further clean sessions where I always started along the lines: "Using this plan {link} do X." Works pretty well for "non-permanent" instructions (you don't want to put this info into your committed markdowns). My main motivation was simply to save tokens, but it actually worked really well and improved speed as well.
- cyanydeez 1mo agowe want some kind of fractal knowledge graph. It starts coarse at some zoom level and you can move in and out. You can insert knowledge at any level and it's out/in levels adjust accordingly. The search semantics at every level are the same, but whats in the visible area changes based on what your focus is. One way I've toyed with a graph outline with it using "whitening" (https://arxiv.org/pdf/2104.01767v3 https://arxiv.org/pdf/2104.01767v3) for embeddings rather than just text, so you add things like file path, nearest title/method/const etc. You have to have dummy text though because it fails with "null"; all embeddings need to carry some kind of text and of the same size. so when the agent rembembers something, the memory would be an embedding that includes where they found the file, what method or const or whatever they're in, what the task they're working on is, etc. That all becomes a single embedding. You could imagine a metadata tag that also describes the tools they're using etc. Map out a complete space of tags for whitening an embedding and there's surely a proper mix. Then when you're searching for things in the embedding, you also store some of the other metadata as plane strings & edges, which gets you some useful granularity.
- wxw 1mo agoHow well does it work in practice?
- tamsinoduya 1mo ago[flagged]
- huksley 1mo agoIt could also be shareable? I just had a thought about similar thing - how to track human decisions on the codebase? Consider you are writing code together with AI, how you understand which code change happened because human asked for it
- Avijit_Thawani 1mo ago"Irrelevant material is simply never surfaced by the semantic search." thats quite optimistic. there's lots of "memory" or past chats with agents that should be suppressed and forgotten because they were looking in the wrong place or were eventually proven wrong. yet semantically they'd look very relevant to a future search. thats why you shouldn't search both textbooks and scifi when trying to solve an examination.
- altruios 1mo agoIt occurs to me: we have latent embedding giving 'general knowledge' to an LLM. What if we use a 'blank' LLM as well as an agent and train that blank LLM on personal context to query that as memory?
- tomashubelbauer 1mo agoCan an LLM be trained to understand language without remembering anything else from its training data? I thought the intrinsic knowledge and the ability to understand language were tied together.
- altruios 1mo agoIt would be closer to using an LLM as a RAG for memory, as the reasoning LLM in injected with the return of the 'memory llm' (maybe with a defined number of 'slots' for easy clean up).
- lksdjfalkjflsdf 1mo agoI think the simplest solution is just directory with md files and https://github.com/BeaconBay/ck https://github.com/BeaconBay/ck
- nsingh2 1mo agoAll of this stuff seems like a band-aid solution. These things need to be trained ground-up to maintain and update persistent memory (maybe outside the context window?). Also seems like a requirement for any sort of continual learning capabilities as well.
- docheinestages 1mo agoMemory is not necessarily a good thing. What we need is a combination of sufficiently large context windows (10-100 million tokens) along with curate, bloat-free data.
- pianopatrick 1mo agoI think eventually you need some kind of system that ranks pieces of data based on how useful they are. I.e. for the web we did that with link count etc. We need some other mechanism for judging and ranking pieces of "memory" for "agents"
- LPisGood 1mo agoAttention is all you need. Seriously, isn’t this the core premise of RAG / modern embedding search systems?
- hedgehog 1mo agoYou can kludge it together by having the agent keep a log of what it does with any notes and lessons about friction / efficiency, and then on some interval review that + session history for items worth promoting into a topic-based memory, or for things that are needed constantly into the main AGENTS.md or a separate MEMORY.md that's loaded into every session. Reviewing the notes at the same time as session history helps provide enough context to make directionally correct decisions about what is worth keeping in context for future sessions.
- smad_ai 1mo ago[flagged]
- BatFastard 1mo ago[flagged]
- JohnMakin 1mo ago> Markdown "pages", with > (optional) YAML frontmatter and > (optional) SQLite vector index for semantic search This is basically exactly what I use in a MCP service I built and it works pretty well. Can be enriched further if you use a storage system like S3 and take advantage of metadata. "harness managed" memory is utter garbage, I am convinced, and I disable it immediately. The major problem being over time it degrades and sneaks in conflicting or outright false information. Then one day you'll swear it's drunk, and every time I got to this state and investigated, auto managed memory was always the problem.
- linggen 1mo agoThe most important thing for a memory system is not only remember and recall or search, it is maintain, that including, merge, forget, update etc. That is how human's work. I am on linggen.dev , that is the one make daily agent work easier.
- loehnsberg 1mo agoThis mirrors my experience: - Store session turns in an sqlite-vec - Provide the agent with an mcp to search the vec-db - Let the agent write notes in md files along with an index / frontmatter Along with the commit history, the vec-db gives the agent long-term memory. The notes allow the user to correct accumulation of false lessons. Simpler but better.
- kelseyfrog 1mo ago> A memoryfield page looks like this: --- title: Carbon Fibre Woks created: '2026-03-01T09:00:00Z' updated: '2026-08-22T14:30:00Z' uuid: 6aa615f0-486f-48a7-a210-ba4f5ff18c8b summary: Thermal properties of carbon fibre cookware --- Carbon fibre woks conduct heat evenly, but... So basically recfiles[1], and an entire suite of tools replaced by fopen and sed. Anyone who doesn't know of recfiles is doomed to reimplement it(poorly). 1. https://www.gnu.org/software/recutils/ https://www.gnu.org/software/recutils/
- calpaterson 1mo agoThanks for pointing out recfiles. I was not aware of it and will investigate whether we can reuse anything. That said Markdown with YAML frontmatter is not exactly my own invention!
- phernandez 1mo ago[flagged]
- hoffiez 1mo ago[flagged]
- _bobm 1mo agoI think this is the wrong approach because everything is external to the model. You end up creating an ad-hoc externalized model scaffolded out of coarser systems, RAGs, files, and so on. This leads to, if useful at all, to this process of ad-hoc recall inference which has to happen in time. By itself this is not a problem. The problem is that the model has to be constantly injected in context with the newest version of the "memory state" at each turn or relevant turn. The newest coherent memory state is also a problem. More or less 50 years of not failures, but not success either. This may be even deeper problem than the externalization problem. I think we are just in the very beginning and we are slapping database stuff to the transformer hoping it will work, but these deep neural-net architectures categorically show that they are not databases. There will be a synergistic middle ground, but its shape is still not clear.
- 2001zhaozhao 1mo agoIf the counterargument to knowledge graph-based memory systems is that they're slow and take multiple steps, then it's not really a counterargument. I'd happily trade off speed for giving the agent ability to find more precise memories. That said, I think this article's basic idea of "text files + semantic search index" is a good way to implement memory because AI is already good at search, and semantic search is far more flexible than a knowledge graph and decreases the chance of the agent simply not being able to find a memory (or inserting duplicate memories, etc.) due to deciding to go into a slightly different branch than the one it needed to travel. That said, I'd probably layer additional steps on top for further efficiency improvements. For example, automatic concise summaries of memory items, other supplemental indexing methods for better memory recall, etc. My ideal solution would be multi-step and therefore wouldn't solve the knowledge-graph's slowness, but it would solve the rigidness problem of a knowledge graph.
- jing09928 1mo agoSeparating declarative facts from active execution skills definitely helps prevent the prompt-drift loop. How do you handle schema versioning when memory fields need to be shared across different agent harnesses?
- calpaterson 1mo agoI'm not sure I understand - the schema is set and doesn't get changed by agents as they work
- mofosyne 1mo agoInstead of having to regenerate a zip file everytime you add/delete a memory... could we just use a git repo of markdown pages? This buys you incremental writes, commit hash pinning and diffs for free. In addition to having a git archive outputting a zip export as well? Main wrinkle is you would need to gitignore the sqlite database as it doesn't store very well in git (binary, changes lots per insert). But it's easy to regenerate anyway as a rebuild-able cache.
- calpaterson 1mo agoYou don't have to regenerate a zipfile each time you do something - the zipfile is just to give a single file for distribution. It's just a zip of the directory. Git is supported in the spec, though the tool doesn't yet handle it (soon! git is nice as you get a log "for free" which helps agents understand more about the memories). As for sqlite in git: yes you can gitignore it, you could use git-lfs, you could use an out of band database. It would be great if there were another format more amenable to git to store vectors in. I looked at csv closely, but I was worried about float formatting/representational bugs. There is a gap in the market for a format here.
- mofosyne 1mo agoAh good to hear! Well regarding how CSV didn't quite work for you. Have you checked out recutils? Its a GNU Project and at least does basic relational database operations on plain text files. You can easily install it in most linux package managers. https://en.wikipedia.org/wiki/Recutils https://en.wikipedia.org/wiki/Recutils
- effnorwood 1mo ago[dead]
- TrustScoreAgent 1mo ago[flagged]
- sangwook 1mo agoThe simplicity of the file format is compelling, but how are page boundaries determined when one memory spans multiple topics? Should related new info update an existing page or crate a new one?
- gimalay 1mo agoThe links traversal is a real N+1 problem, but it doesn't need to be performed by the agent. A tool can pre-assemble all the relevant context by following the links in deterministic way. https://iwe.md/blog/your-agent-hates-walking-your-knowledge-graph/ https://iwe.md/blog/your-agent-hates-walking-your-knowledge-...
- r14c 1mo agoIt isn't done enough to share yet, but I've been having fun experimenting with a knowledge graph backed by sqlite, using the tool call schema from mcp-server-memory. My main gripe so far is that I have to push the agent to maintain an organized graph. Once its there, it can be pretty useful to keep track of small notes which would otherwise be sprinkles all over my projects. The kg has some design limits that make poisoning hard to diagnose, so that's still on the list to fix.
- leouno 1mo ago[flagged]
- soricus 1mo agoNegative memory help but only if someone checks it with a separate pass. The presence of the entry "this is incorrect" in the context does not guarantee anything.
- beyondscale-sai 1mo ago[dead]